簡易檢索 / 詳目顯示

研究生: 林辰安
Lin, Chen-An
論文名稱: 一個用於單影像耀光去除之高效率頻率分解混和式網路
An Efficient Frequency-Decomposed Hybrid Network for Single Image Flare Removal
指導教授: 戴顯權
Tai, Shen-Chuan
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 66
中文關鍵詞: 耀光去除拉普拉斯金字塔深度學習影像復原
外文關鍵詞: flare removal, Laplacian pyramid, deep learning, image restoration
相關次數: 點閱:4下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 鏡頭拍攝影像時,尤其夜間攝影中,路燈、車燈與招牌等高亮度光源容易在鏡片組、保護鏡與感光元件之間產生多次反射,形成鬼影、光斑與條紋等反射耀光;同時,鏡片表面缺陷、髒污或介質散射亦會造成霧狀光暈與低對比區域。這些散射與反射耀光會遮蔽物體邊界、破壞紋理細節,進而降低自動駕駛、夜間監控與智慧交通等電腦視覺任務之可靠性。
    近年來,Transformer 架構因具備長距離相依關係建模能力,已廣泛應用於影像復原;其中 Uformer 透過階層式編解碼器與自注意力機制,可有效提升去耀光品質。然而,自注意力在高解析度影像上仍伴隨龐大參數量、運算量與記憶體需求,使其難以部署於邊緣裝置、車載系統或即時夜間影像應用。
    為解決上述問題,本論文提出基於拉普拉斯金字塔的輕量化 Uformer。先將影像分解為低頻與高頻成分,低頻影像由 Uformer 擷取全域照明與結構特徵,高頻殘差則交由輕量 CNN 分支與 FNetBlock 進行全域混合與細節修正。實驗結果顯示,此方法在維持良好夜間耀光去除品質的同時,大幅降低硬體開銷,說明頻率分解與輕量高頻分支有助於兼顧復原效能與部署效率。

    When capturing images, especially in nighttime photography, strong light sources such as street lamps, vehicle headlights, and illuminated signs can easily cause multiple reflections among lens elements, protective glass, and the image sensor, producing reflective flare artifacts such as ghosting, bright spots, and streaks. Meanwhile, lens surface defects, contamination, or medium scattering may also lead to veiling glare and low-contrast regions. These scattering and reflective flares obscure object boundaries and damage texture details, thereby reducing the reliability of computer vision tasks such as autonomous driving, nighttime surveillance, and intelligent transportation.
    In recent years, Transformer architectures have been widely applied to image restoration due to their ability to model long-range dependencies. Among them, Uformer can effectively improve flare removal quality through a hierarchical encoder--decoder structure and self-attention mechanism. However, self-attention still introduces large parameter counts, high computational cost, and substantial memory requirements when processing high-resolution images, making it difficult to deploy on edge devices, in-vehicle systems, or real-time imaging applications.
    To address the above problem, this Thesis proposes a Laplacian Pyramid-Enhanced Uformer. The proposed method first decomposes the image into low-frequency and high-frequency components. The low-frequency image is processed by a three-level Uformer backbone consisting of encoder, bottleneck, and decoder modules to extract global illumination and structural features, while the high-frequency residuals are handled by a lightweight CNN branch and a two-dimensional fast Fourier transform-based FNetBlock for global mixing and detail refinement. Experimental results show that the proposed method significantly reduces computational cost while maintaining good nighttime flare removal quality, demonstrating that frequency decomposition, low-frequency Transformer modeling, and a lightweight high-frequency branch can effectively balance restoration performance and deployment efficiency.

    中文摘要 i Abstract ii Acknowledgements iv Contents v List of Tables viii List of Figures ix Chapter 1 Introduction 1 Chapter 2 Related Works 5 2.1 Nighttime Flare Removal 5 2.1.1 Traditional Flare Removal 5 2.1.2 Deep Learning-based Flare Removal 6 2.2 Transformer-based Image Restoration 8 2.2.1 Vision Transformer 9 2.2.2 Swin Transformer 9 2.2.3 Cross-Attention 10 2.2.4 Uformer 10 2.3 Frequency-aware Image Restoration 12 2.3.1 Frequency Decomposition Networks 12 2.3.2 SAFAFormer 13 2.3.3 Laplacian Pyramid 14 2.4 Fourier-based Feature Mixing 15 2.4.1 Fourier Transform in Vision 15 2.4.2 FNet 16 2.5 FBNet 17 2.6 Discussion and Motivation 19 Chapter 3 The Proposed Method 21 3.1 Overall Architecture 22 3.2 Uformer Backbone 26 3.3 Lightweight High-Frequency Branch 28 3.4 Double-head Output Disentanglement 31 3.5 Loss Function 32 3.5.1 Absolute Error Loss 33 3.5.2 Perceptual Loss 34 3.5.3 Physical Reconstruction Loss 34 3.5.4 LPIPS Loss 35 3.5.5 Total Loss 35 Chapter 4 Performance Evaluation 37 4.1 Experimental Dataset 37 4.1.1 Training Dataset 37 4.1.2 Testing Dataset 38 4.2 Experimental Settings 39 4.2.1 Training Hyperparameters 39 4.2.2 Implementation Environment 40 4.3 Experimental Results 40 4.3.1 Evaluation Metrics 40 4.3.2 Quantitative Results 43 4.3.3 Visual Results Comparison 44 4.4 Ablation Experimental Results 46 Chapter 5 Conclusions 48 5.1 Conclusions 48 5.2 Future Work 49 References 51

    [1] J. Lian, J. Liu, M. Jing, X. Zeng, Z. Liu, J. Zhou, and Y. Fan, “Nighttime Glare Removal for Consumer Electronics via Latent Space Transformation and Feature-Enhanced Attention Mechanism,” IEEE Transactions on Consumer Electronics, vol. 71, no. 2, pp. 6719–6733, 2025.
    [2] Y. Dai, C. Li, S. Zhou, R. Feng, Y. Luo, and C. C. Loy, “Flare7K++: Mixing Synthetic and Real Datasets for Nighttime Flare Removal and Beyond,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 11, pp. 7041–7055, 2024.
    [3] J. Lee-Thorp, J. Ainslie, I. Eckstein, and S. Ontanon, “FNet: Mixing Tokens with Fourier Transforms,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 4296–4313, 2022.
    [4] O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional Networks for Biomedical Image Segmentation,” in Medical Image Computing and Computer-Assisted Intervention. Springer, pp. 234–241, 2015.
    [5] Y. Wu, Q. He, T. Xue, R. Garg, J. Chen, A. Veeraraghavan, and J. T. Barron, “How to Train Neural Networks for Flare Removal,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2239–2247, 2021.
    [6] Y. Dai, C. Li, S. Zhou, R. Feng, and C. C. Loy, “Flare7K: A Phenomenological Nighttime Flare Removal Dataset,” in Advances in Neural Information Processing Systems Datasets and Benchmarks Track, 2022.
    [7] Y. He, W. Wang, W. Wu, and K. Jiang, “Disentangle Nighttime Lens Flares: SelfSupervised Generation-Based Lens Flare Removal,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 3, pp. 3464–3472, 2025.
    [8] K. Qi, B. Wang, and C. Li, “FCNet: A Feature Complementary Network for Nighttime Flare Removal,” Computer Vision and Image Understanding, vol. 261, p. 104495, 2025.
    [9] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in International Conference on Learning Representations, 2021.
    [10] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10 012–10 022, 2021.
    [11] D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in International Conference on Learning Representations, 2015.
    [12] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in Advances in Neural Information Processing Systems, vol. 30, 2017.
    [13] Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li, “Uformer: A General U-Shaped Transformer for Image Restoration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 17 683–17 693, 2022.
    [14] S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M.-H. Yang, “Restormer: Efficient Transformer for High-Resolution Image Restoration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5728–5739, 2022.
    [15] S. Wu, F. Liu, Y. Bai, H. Han, J. Wang, and N. Zhang, “Flare Removal Model Based on Sparse-UFormer Networks,” Entropy, vol. 26, no. 8, p. 627, 2024.
    [16] Y. Chen, H. Fan, B. Xu, Z. Yan, Y. Kalantidis, M. Rohrbach, S. Yan, and J. Feng, “Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks With Octave Convolution,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3435–3444, 2019.
    [17] S. Liu, G. Chen, W. Zhang, Q. Zhang, J. Zhang, and J. Wang, “SMFR-Net: Simple Multi-Domain Flare Removal Network,” Scientific Reports, vol. 15, p. 37251, 2025.
    [18] W. Dong, G. Fan, F. Zhang, M. Gan, G.-Y. Chen, and C. L. P. Chen, “SAFAFormer: Scale-Aware Frequency-Adaptive Guidance for Nighttime Flare Removal,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 36, no. 1, pp. 93–105, 2026.
    [19] P. J. Burt and E. H. Adelson, “The Laplacian Pyramid as a Compact Image Code,” IEEE Transactions on Communications, vol. 31, no. 4, pp. 532–540, 1983.
    [20] W.-S. Lai, J.-B. Huang, N. Ahuja, and M.-H. Yang, “Deep Laplacian Pyramid Networks for Fast and Accurate Super-Resolution,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 624–632, 2017.
    [21] L. Kong, J. Dong, J. Ge, M. Li, and J. Pan, “Efficient Frequency Domain-Based Transformers for High-Quality Image Deblurring,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5886–5895, 2023.
    [22] D. Zhang, J. Ouyang, G. Liu, X. Wang, X. Kong, and Z. Jin, “FF-Former: Swin Fourier Transformer for Nighttime Flare Removal,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 2824–2832, 2023.
    [23] J. Ren, Z. Zhang, S. Zhao, J. Fan, Z. Zhao, Y. Zhao, R. Hong, and M. Wang, “When Low-Light Meets Flares: Towards Synchronous Flare Removal and Brightness Enhancement,” Neural Networks, vol. 185, p. 107149, 2025.
    [24] B. Yang, “Projection Approximation Subspace Tracking,” IEEE Transactions on Signal Processing, vol. 43, no. 1, pp. 95–107, Jan. 1995.
    [25] A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier Nonlinearities Improve Neural Network Acoustic Models,” in Proceedings of the ICML Workshop on Deep Learning for Audio, Speech and Language Processing, 2013.
    [26] Y. Jiang, X. Chen, C.-M. Pun, S. Wang, and W. Feng, “MFDNet: Multi-Frequency Deflare Network for Efficient Nighttime Flare Removal,” The Visual Computer, vol. 40, no. 11, pp. 7575–7588, 2024.
    [27] K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” in International Conference on Learning Representations, 2015.
    [28] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 586–595, 2018.
    [29] O. Lee, “Flickr24K,” Kaggle Datasets. [Online]. Available: https://www.kaggle.com/datasets/okitalee/flickr24k Accessed: Jan. 15, 2026, 2022.
    [30] D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” in International Conference on Learning Representations, 2015.
    [31] Q. Huynh-Thu and M. Ghanbari, “Scope of Validity of PSNR in Image/Video Quality Assessment,” Electronics Letters, vol. 44, no. 13, pp. 800–801, 2008.
    [32] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image Quality Assessment: From Error Visibility to Structural Similarity,” IEEE Transactions on Image Processing, vol. 13, no. 4, pp. 600–612, 2004.
    [33] T. Ma, Z. Kai, X. Miao, J. Liang, J. Peng, Y. Wang, H. Wang, and X. Liu, “SelfPrior Guided Spatial and Fourier Transformer for Nighttime Flare Removal,” IEEE Transactions on Automation Science and Engineering, vol. 22, pp. 11 996–12 011, 2025.

    QR CODE