簡易檢索 / 詳目顯示

研究生: 林宇凡
Lin, Yu-Fan
論文名稱: 基於場景先驗引導之單張影像陰影移除
PriorSR: Prior-Guided Single-Image Shadow Removal
指導教授: 許志仲
Hsu, Chih-Chung
學位類別: 碩士
Master
系所名稱: 敏求智慧運算學院 - 智慧科技系統碩士學位學程
MS Degree Program on Intelligent Technology Systems
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 50
中文關鍵詞: 單張影像陰影移除場景先驗多模態學習注意力機制
外文關鍵詞: Single-Image Shadow Removal, Scene Priors, Multi-modality Learning, Attention Mechanism
相關次數: 點閱:1下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究提出一種名為 PriorSR 的單張影像陰影移除框架,透過場景先驗的引導,以修復品質為核心目標。單張影像陰影移除的困難在於陰影造成的退化在空間上並不均勻,且暗區可能來自光線遭到遮蔽,也可能來自物體本身的深色材質,兩者難以區分。此困難在間接光照下更為明顯,因為此時柔和且漸層的陰影佔多數,傳統方法往往被迫在兩個目標之間取捨,不是抹去陰影內部的紋理,就是模糊其邊界。

    為此,PriorSR 採取「先理解、後修復」的分層策略。在理解階段,先驗調變注意力模組(PMAM)以注意力機制引導幾何先驗(深度、法向量)與語意先驗(DINO 特徵),達成場景理解與陰影消歧。在修復階段,解碼器中創新的適應性融合區塊(AFB)對特徵成分進行選擇性處理,由內容一致性模組(CCM)重建陰影內部均勻的基底外觀,並由細節精煉模組(DRM)復原精細紋理與銳利邊界,融合兩股互補的特徵流後,即可兼顧內部一致性與細節豐富度。

    在 ISTD、ISTD+、SRD、WSRD+ 與 INS 資料集上的實驗顯示,PriorSR 於所有基準的 SSIM 均為最佳,並在多數資料集上取得最佳 PSNR,在複雜的直接與間接光照場景中展現穩定的優勢。

    Shadows are among the most common artifacts that compromise image quality. Recovering the true appearance beneath a shadow is intrinsically difficult, since the degradation is spatially uneven and a darkened region may stem from either occluded light or the object's own dark material. The difficulty grows under indirect illumination, where soft, graded shadows dominate. Conventional approaches therefore trade one goal off against another, either washing out the texture inside the shadow or smearing its boundaries. This thesis places restoration fidelity at the center of the task and proposes the PriorSR framework around two coordinated ideas. A scene-understanding stage injects geometric and semantic priors through attention to disambiguate and implicitly localize shadows. A restoration stage then builds on a novel Adaptive Fusion Block (AFB) in the decoder, where a Content Consistency Module (CCM) reconstructs a uniform base appearance within shadowed regions and a Detail Refinement Module (DRM) recovers fine textures and re-sharpens boundaries. Fusing the two complementary streams yields features that are at once internally consistent and rich in detail. Experiments on ISTD, ISTD+, SRD, WSRD+, and INS show that PriorSR leads in SSIM on every benchmark and achieves the best PSNR on most, with consistent advantages on scenes governed by complex direct and indirect lighting.

    中文摘要 i Abstract ii Acknowledgements iv Contents v List of Tables vii List of Figures viii 1 Introduction 1 1.1 Background 1 1.2 Motivation 2 1.3 Overview of the Proposed Method 3 1.4 Contributions 4 1.5 Thesis Organization 5 2 Related Works 6 2.1 Single-Image Shadow Removal 6 2.2 Dense Prediction 7 2.3 Vision Foundation Models and Prior Knowledge 8 3 Methodology 10 3.1 Shadow Physics, Image Model, and Challenges 10 3.2 Evolution of Learning Strategies and the Motivation for PriorSR 11 3.3 Network Architecture Overview 12 3.3.1 Spatial and Scene-Prior Extraction and Preprocessing 14 3.3.2 Prior-Modulated Attention Module (PMAM) 14 3.4 Adaptive Fusion for Shadow Removal 17 3.4.1 Content Consistency Module (CCM) 17 3.4.2 Detail Refinement Module (DRM) 18 3.4.3 Component Integration within the AFB 19 3.5 Training Objective 20 4 Experimental Results 21 4.1 Implementation Details 21 4.2 Performance Comparisons 22 4.3 Ablation Study 24 5 Conclusions and Future Work 29 5.1 Conclusion 29 5.2 Future Work 30 References 31

    [1] L. Bolanos, S.-Y. Su, and H. Rhodin, “Gaussian shadow casting for neural characters,” in The Conference on Computer Vision and Pattern Recognition, 2024.
    [2] N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in European conference on computer vision. Springer, 2020, pp. 213–229.
    [3] L. Chen, Y. Fu, L. Gu, C. Yan, T. Harada, and G. Huang, “Frequency-aware feature fusion for dense image prediction,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.
    [4] R. Cucchiara, C. Grana, M. Piccardi, and A. Prati, “Detecting moving objects, ghosts, and shadows in video streams,” IEEE transactions on pattern analysis and machine intelligence, vol. 25, no. 10, pp. 1337–1342, 2003.
    [5] X. Cun, C.-M. Pun, and C. Shi, “Towards ghost-free shadow removal via dual hierarchical aggregation network and shadow matting gan.” 2020, pp. 10680–10687.
    [6] T. Daboczi and T. Bako, “Inverse filtering of optical images,” in Proceedings of the 17th IEEE Instrumentation and Measurement Technology Conference [Cat. No. 00CH37066], vol. 1, 2000, pp. 370–374 vol.1.
    [7] R. K. Das, M. Shandilya, S. Sharma, and D. Kulkarni, “A survey on shadow detection and removal in images,” in 2017 International Conference on Recent Innovations in Signal processing and Embedded Systems (RISE), 2017, pp. 175–180.
    [8] W. Dong, H. Zhou, Y. Tian, J. Sun, X. Liu, G. Zhai, and J. Chen, “Shadowrefiner: Towards mask-free shadow removal via fast fourier transformer,” 2024. [Online]. Available: https://arxiv.org/abs/2406.02559
    [9] L. Fu, C. Zhou, Q. Guo, F. Juefei-Xu, H. Yu, W. Feng, Y. Liu, and S. Wang, “Auto-exposure fusion for single-image shadow removal,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 10571–10580.
    [10] A. A. Goshtasby, “Fusion of multi-exposure images,” Image and Vision Computing, vol. 23, no. 6, pp. 611–618, 2005.
    [11] L. Guo, C. Wang, Y. Wang, S. Huang, W. Yang, A. C. Kot, and B. Wen, “Single-image shadow removal using deep learning: A comprehensive survey,” arXiv preprint arXiv:2407.08865, 2024.
    [12] L. Guo, S. Huang, D. Liu, H. Cheng, and B. Wen, “Shadowformer: global context helps shadow removal,” in Proceedings of the AAAI conference on artificial intelligence, vol. 37, no. 1, 2023, pp. 710–718.
    [13] L. Guo, C. Wang, W. Yang, S. Huang, Y. Wang, H. Pfister, and B. Wen, “Shadowdiffusion: When degradation prior meets diffusion model for shadow removal,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 14049–14058.
    [14] B. Gwangbin, B. Ignas, and C. Roberto, “Estimating and exploiting the aleatoric uncertainty in surface normal estimation,” in International Conference on Computer Vision (ICCV), 2021.
    [15] K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in 2017 IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2980–2988.
    [16] S. He, B. Peng, J. Dong, and Y. Du, “Mask-shadownet: Toward shadow removal via masked adaptive instance normalization,” IEEE Signal Processing Letters, vol. 28, pp. 957–961, 2021.
    [17] C.-C. Hsu, C.-Y. Jian, E.-S. Tu, C.-M. Lee, and G.-L. Chen, “Real-time compressed sensing for joint hyperspectral image transmission and restoration for cubesat,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–16, 2024.
    [18] C.-C. Hsu, C.-M. Lee, and Y.-S. Chou, “Drct: Saving image super-resolution away from information bottleneck,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2024, pp. 6133–6142.
    [19] C.-C. Hsu, C.-M. Lee, Y.-F. Lin, Y.-S. Chou, C.-Y. Jian, and C.-H. Tsai, “Revisiting vision-language features adaptation and inconsistency for social media popularity prediction,” in Proceedings of the 32nd ACM International Conference on Multimedia, ser. MM ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 11464–11469. [Online]. Available: https://doi.org/10.1145/3664647.3689000
    [20] X. Hu, C.-W. Fu, L. Zhu, J. Qin, and P.-A. Heng, “Direction-aware spatial context features for shadow detection and removal,” IEEE transactions on pattern analysis and machine intelligence, vol. 42, no. 11, pp. 2795–2808, 2019.
    [21] Z. Huang, Y. Wei, X. Wang, W. Liu, T. S. Huang, and H. Shi, “Alignseg: Feature-aligned segmentation networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 1, pp. 550–557, 2022.
    [22] H. Le and D. Samaras, “Shadow removal via shadow image decomposition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 8578–8587.
    [23] ——, “From shadow segmentation to shadow removal,” 2020.
    [24] C.-M. Lee, C.-H. Cheng, Y.-F. Lin, Y.-C. Cheng, W.-T. Liao, F.-E. Yang, Y.-C. F. Wang, and C.-C. Hsu, “Prompthsi: Universal hyperspectral image restoration with vision-language modulated frequency adaptation,” 2025. [Online]. Available: https://arxiv.org/abs/2411.15922
    [25] Y. Lee, J. Kim, J. Willette, and S. J. Hwang, “Mpvit: Multi-path vision transformer for dense prediction,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 7277–7286.
    [26] C. Li, B. Yang, Z. Wu, G. Chen, Y. Yu, and S. Zhou, “Shadow removal based on diffusion, segmentation and super-resolution models,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2024, pp. 6045–6054.
    [27] T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 936–944.
    [28] H. Liu, M. Li, and X. Guo, “Regional attention for shadow removal,” 2024. [Online]. Available: https://arxiv.org/abs/2411.14201
    [29] J. Liu, Q. Wang, H. Fan, W. Li, L. Qu, and Y. Tang, “A decoupled multi-task network for shadow removal,” IEEE Transactions on Multimedia, 2023.
    [30] J. Liu, Q. Wang, H. Fan, J. Tian, and Y. Tang, “A shadow imaging bilinear model and three-branch residual network for shadow removal,” IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 11, pp. 15857–15871, 2024.
    [31] W. Liu, H. Lu, H. Fu, and Z. Cao, “Learning to upsample by learning to sample,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 6004–6014.
    [32] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10012–10022.
    [33] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” 2019. [Online]. Available: https://arxiv.org/abs/1711.05101
    [34] H. Lu, W. Liu, Z. Ye, H. Fu, Y. Liu, and Z. Cao, “Sapa: Similarity-aware point affiliation for feature upsampling,” in Proc. Annual Conference on Neural Information Processing Systems (NeurIPS), 2022.
    [35] K. Mei, L. Figueroa, Z. Lin, Z. Ding, S. Cohen, and V. M. Patel, “Latent feature-guided diffusion models for shadow removal,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 4313–4322.
    [36] S. Murali, V. Govindan, and S. Kalady, “A survey on shadow removal techniques for single image,” International Journal of Image, Graphics and Signal Processing, vol. 8, no. 12, p. 38, 2016.
    [37] K. Niu, Y. Liu, E. Wu, and G. Xing, “A boundary-aware network for shadow removal,” IEEE Transactions on Multimedia, vol. 25, pp. 6782–6793, 2023.
    [38] M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski, “Dinov2: Learning robust visual features without supervision,” 2024. [Online]. Available: https://arxiv.org/abs/2304.07193
    [39] L. Qu, J. Tian, S. He, Y. Tang, and R. W. Lau, “Deshadownet: A multi-context embedding deep network for shadow removal,” 2017, pp. 4067–4075.
    [40] R. Ranftl, A. Bochkovskiy, and V. Koltun, “Vision transformers for dense prediction,” in 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 12159–12168.
    [41] O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention. Springer, 2015, pp. 234–241.
    [42] N. Salamati, A. Germain, and S. Süsstrunk, “Removing shadows from images using color and near-infrared,” in 2011 18th IEEE International Conference on Image Processing. IEEE, 2011, pp. 1713–1716.
    [43] A. Sanin, C. Sanderson, and B. C. Lovell, “Improved shadow removal for robust person tracking in surveillance scenarios,” in 2010 20th International Conference on Pattern Recognition, 2010, pp. 141–144.
    [44] P. Sharma, J. Philip, M. Gharbi, B. Freeman, F. Durand, and V. Deschaintre, “Materialistic: Selecting similar materials in images,” ACM Trans. Graph., vol. 42, no. 4, jul 2023. [Online]. Available: https://doi.org/10.1145/3592390
    [45] E. Shelhamer, J. Long, and T. Darrell, “Fully convolutional networks for semantic segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 4, pp. 640–651, 2017.
    [46] W. Shi, J. Caballero, F. Huszár, J. Totz, A. P. Aitken, R. Bishop, D. Rueckert, and Z. Wang, “Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 1874–1883.
    [47] Y. Shor and D. Lischinski, “The shadow meets the mask: Pyramid-based shadow removal,” in Computer Graphics Forum, vol. 27, no. 2. Wiley Online Library, 2008, pp. 577–586.
    [48] A. Tiwari, P. K. Singh, and S. Amin, “A survey on shadow detection and removal in images and video sequences,” in 2016 6th International Conference-Cloud System and Big Data Engineering (Confluence). IEEE, 2016, pp. 518–523.
    [49] F.-A. Vasluianu, T. Seizinger, and R. Timofte, “Wsrd: A novel benchmark for high resolution image shadow removal,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2023, pp. 1826–1835.
    [50] F.-A. Vasluianu, T. Seizinger, Z. Zhou, Z. Wu, C. Chen, and R. Timofte, “Ntire 2024 image shadow removal challenge report,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2024, pp. 6547–6570.
    [51] J. Wang, K. Chen, R. Xu, Z. Liu, C. C. Loy, and D. Lin, “Carafe: Content-aware reassembly of features,” in The IEEE International Conference on Computer Vision (ICCV), October 2019.
    [52] J. Wang, X. Li, and J. Yang, “Stacked conditional generative adversarial networks for jointly learning shadow detection and shadow removal,” 2018, pp. 1788–1797.
    [53] W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 568–578.
    [54] Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li, “Uformer: A general u-shaped transformer for image restoration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 17683–17693.
    [55] S. Weder, G. Garcia-Hernando, A. Monszpart, M. Pollefeys, G. J. Brostow, M. Firman, and S. Vicente, “Removing objects from neural radiance fields,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 16528–16538.
    [56] J. Xiao, X. Fu, Y. Zhu, D. Li, J. Huang, K. Zhu, and Z.-J. Zha, “Homoformer: Homogenized transformer for image shadow removal,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024, pp. 25617–25626.
    [57] J. Xu, Z. Li, Y. Zheng, C. Huang, R. Gu, W. Xu, and G. Xu, “Omnisr: Shadow removal under direct and indirect lighting,” 2025. [Online]. Available: https://arxiv.org/abs/2410.01719
    [58] J. Xu, Y. Zheng, Z. Li, C. Wang, R. Gu, W. Xu, and G. Xu, “Detail-preserving latent diffusion for stable shadow removal,” 2024. [Online]. Available: https://arxiv.org/abs/2412.17630
    [59] L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” 2024. [Online]. Available: https://arxiv.org/abs/2406.09414
    [60] Q. Yang, K.-H. Tan, and N. Ahuja, “Shadow removal using bilateral filtering,” IEEE Transactions on Image processing, vol. 21, no. 10, pp. 4361–4368, 2012.
    [61] S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, M.-H. Yang, and L. Shao, “Learning enriched features for real image restoration and enhancement,” in IEEE/CVF European Conference on Computer Vision (ECCV), 2020.
    [62] L. Zhang, Q. Zhang, and C. Xiao, “Shadow remover: Image shadow removal based on illumination recovering optimization,” IEEE Transactions on Image Processing, vol. 24, no. 11, pp. 4623–4636, 2015.
    [63] Y. Zhu, J. Huang, X. Fu, F. Zhao, Q. Sun, and Z.-J. Zha, “Bijective mapping network for shadow removal,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5627–5636.
    [64] Y. Zhu, Z. Xiao, Y. Fang, X. Fu, Z. Xiong, and Z.-J. Zha, “Efficient model-driven network for shadow removal,” in Proceedings of the AAAI conference on artificial intelligence, vol. 36, no. 3, 2022, pp. 3635–3643.

    QR CODE