簡易檢索 / 詳目顯示

研究生: 呂順瑋
LU, Shun-Wei
論文名稱: 一個基於低頻特徵門控的高效全能影像修復網路
An Efficient Low-Frequency Feature Gating Network for All-in-One Image Restoration
指導教授: 戴顯權
Tai, Shen-Chuan
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 102
中文關鍵詞: 全能影像修復低頻特徵門控退化感知門控門控前饋網路
外文關鍵詞: All-in-One Image Restoration, Low-Frequency Feature Gating, Degradation-Aware Gating, Gated Feed-Forward Network
相關次數: 點閱:13下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 全能影像修復(All-in-One Image Restoration, AiOIR)旨在以單一模型修復受多種未知退化影響的影像,避免為每種退化各別維護一組權重所衍生、隨退化種類數線性增長的儲存與存取需求,以及實際場景中退化難以預知、卻仍須預先選定對應模型的限制。在各類 AiOIR 架構中,不依賴外部退化提示的統一框架具有最高的計算效率,也最適合部署於資源受限的環境。近期提出的 AnyIR 即為此一範式樹立了強而精簡的基準:僅以約 5.74M 參數,便達到與數倍規模之模型(如基於提示的 PromptIR)相當甚至更優的修復品質,位居輕量級 AiOIR 中品質與效率的權衡前緣。然而,此類模型雖已逼近效率前緣,其部分設計仍沿襲既有慣例,未必為當前情境之最優。在不擴增模型容量的前提下重新檢視此等設計,即為本研究尋求進一步效能增益之切入點。

    基於上述觀察,本研究以 AnyIR 為基準,提出兩項互補且機制獨立的改進,並整合為低頻特徵門控網路(Low-Frequency Feature Gating Network, LFGNet)。其一為低頻特徵門控(LFG):基於退化能量主要集中於低頻成分之頻域證據,將退化感知門控所用之全解析度統計量替換為低頻成分統計量,以排除與退化無關之高頻細節干擾,其低頻統計量以 2 * 2 平均池化取得,在常數倍率的意義下等價於擷取 Haar 小波轉換之低頻子帶,且不增加任何參數。其二為潛在層前饋網路(latent FFN)展開率調整:基準模型所沿用之前饋網路在門控激活下實為有效四倍展開,高於門控前饋網路之理論建議值,遂將位於網路最深處、參數佔比最大之潛在層展開率調降至該建議值,在縮減模型參數的同時亦小幅降低計算量。

    在 All-in-One 三種退化的標準評估下,針對兩項改動之有無進行全部四種組合的消融實驗,證實兩者機制獨立、效應大致可分離,僅存在集中於去雨之小幅相互抵銷。兩項改動遂將基準的單一設計點推進為品質與效率權衡上的兩個可選配置:品質優先之 LFGNet-B 於與基準模型相同的參數量下,平均 PSNR 較復現之基準提升 0.10 dB,甚至略優於基準模型之變體 AnyIR-S,後者參數量多出將近一半;效率優先之 LFGNet-T 則將參數量減少 12.7%,修復品質仍略優於基準。

    All-in-One Image Restoration (AiOIR) aims to restore images affected by multiple unknown degradations with a single model. This removes the storage and access costs of per-degradation weight sets, which grow linearly with the number of types, and the need to select a matching model in advance despite the difficulty of predicting real-world degradations. Among AiOIR designs, the pure unified framework that does not rely on external degradation prompts is the most computationally efficient and best suited to resource-constrained deployment. The recent AnyIR establishes a strong yet compact baseline for this paradigm: with only about 5.74M parameters, it attains restoration quality comparable to, or better than, that of models several times larger, such as the prompt-based PromptIR, placing it on the quality--efficiency frontier of lightweight AiOIR. Nevertheless, some design choices in such models still follow established conventions and are not necessarily optimal for the present setting. Re-examining these choices without enlarging model capacity is the entry point from which this Thesis seeks further performance gains.

    Building on this observation, this Thesis takes AnyIR as the baseline, proposes two complementary, mechanistically independent improvements, and integrates them into the Low-Frequency Feature Gating Network (LFGNet). The first, termed Low-Frequency Feature Gating (LFG), is motivated by frequency-domain evidence that degradation energy is concentrated primarily in low-frequency components. It replaces the full-resolution statistics used in degradation-aware gating with low-frequency statistics, excluding interference from high-frequency details unrelated to degradation. These statistics are obtained by 2 * 2 average pooling, which is equivalent up to a constant scale to extracting the Haar wavelet low-frequency sub-band and introduces no additional parameters. The second is a latent FFN expansion-ratio adjustment. Under its gated activation, the feed-forward network inherited by the baseline effectively has an expansion ratio of four, above the theoretically recommended value for gated FFNs. The expansion ratio of the latent stage, the network's deepest and most parameter-intensive part, is therefore adjusted to that value, which reduces the model size and slightly lowers computational cost.

    Under the standard All-in-One evaluation with three degradation types, an ablation over all four combinations confirms that the two improvements are mechanistically independent and their effects largely separable. Only a small offsetting interaction is observed, concentrated in the deraining task. They thus transform the baseline's single design point into two selectable configurations on the quality--efficiency trade-off. At the baseline's parameter budget, the quality-oriented LFGNet-B improves average PSNR by 0.10 dB over the reproduced baseline and even slightly surpasses AnyIR-S, a larger variant with nearly 50% more parameters. In contrast, the efficiency-oriented LFGNet-T reduces the parameter count by 12.7% while maintaining restoration quality slightly above the baseline.

    中文摘要 i Abstract iii Acknowledgements v Contents vii List of Tables xi List of Figures xii List of Symbols xiii Chapter 1 Introduction 1 1.1 Background and Motivation 1 1.2 Research Problems and Challenges 3 1.3 Methods and Contributions 5 1.4 Thesis Organization 7 Chapter 2 Background and Related Work 8 2.1 Development of Restoration Networks 8 2.2 All-in-One Paradigms and Baseline Selection 10 2.3 Frequency-Domain Basis of Degradation Awareness 12 2.4 Efficient Transformers and Cost Distribution 15 2.5 Expansion Ratio of Gated FFNs 16 Chapter 3 The Proposed Method 18 3.1 Overall Architecture of LFGNet 18 3.1.1 Degradation-Aware Block 20 3.1.2 Gated Degradation Adaption (GatedDA) 22 3.1.3 Spatial-Frequency Fusion (SFF) 23 3.1.4 Feed-Forward Network 24 3.2 Low-Frequency Feature Gating (LFG) 25 3.2.1 Low-Frequency Statistics as the Degradation Signal 27 3.2.2 Equivalence of the Haar LL Sub-band and Average Pooling 28 3.2.3 Expected Effect on Each Degradation Task 30 3.3 Latent FFN Expansion-Ratio Adjustment 32 3.3.1 Analysis of the GEGLU Effective Expansion Ratio 33 3.3.2 Latent-Only Adjustment of the Expansion Ratio 33 3.4 Combining the Two Changes and Model Configurations 34 3.4.1 Mechanistic Independence and Complementarity 34 3.4.2 Model Configurations: LFGNet-B and LFGNet-T 35 3.5 Loss Function 37 Chapter 4 Performance Evaluation 38 4.1 Experimental Datasets 38 4.1.1 Training Datasets 38 4.1.2 Testing Datasets 39 4.2 Experimental Settings 39 4.2.1 Training Protocol 39 4.2.2 FLOPs Accounting Protocol 40 4.2.3 Hardware Environment 41 4.3 Evaluation Metrics 41 4.3.1 Peak Signal-to-Noise Ratio (PSNR) 42 4.3.2 Structural Similarity Index Measure (SSIM) 42 4.4 Restoration Quality Comparison 43 4.4.1 Baseline Reproduction Validation 43 4.4.2 Comparison with State-of-the-Art Methods 44 4.4.3 Visual Comparison 47 4.5 Efficiency Analysis 57 4.5.1 Complexity Comparison 57 4.5.2 Inference Latency and Memory 58 4.6 Ablation Study 60 4.7 Discussion 62 4.7.1 Per-Task Analysis 63 4.7.2 Per-Image Analysis of Deraining 64 4.7.3 Failure Cases 70 Chapter 5 Conclusions and Future Work 72 5.1 Conclusions 72 5.1.1 Research Contributions 72 5.1.2 Main Empirical Results 73 5.1.3 Significance of the Research 74 5.2 Limitations 75 5.3 Future Work 76 References 79

    [1] B. Ren, E. Zamfir, Z. Wu, Y. Li, Y. Li, D. P. Paudel, R. Timofte, M.-H. Yang, L. Van Gool, and N. Sebe, “Any image restoration via efficient spatial-frequency degradation adaptation,” Transactions on Machine Learning Research (TMLR), 2026.
    [2] V. Potlapalli, S. W. Zamir, S. Khan, and F. S. Khan, “PromptIR: Prompting for all-in-one blind image restoration,” in Advances in Neural Information Processing Systems (NeurIPS), 2023.
    [3] B. Wang, B. Qin, J. Li, F. Xu, F. Sun, and H. Xiong, “All-in-one image restoration via causal-deconfounding wavelet-disentangled prompt network,” IEEE Transactions on Image Processing, vol. 35, pp. 3421–3436, 2026.
    [4] S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M.-H. Yang, “Restormer: Efficient transformer for high-resolution image restoration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5718–5729, 2022.
    [5] N. Shazeer, “GLU variants improve transformer,” arXiv preprint arXiv:2002.05202, 2020.
    [6] D. Hendrycks and K. Gimpel, “Gaussian error linear units (GELUs),” arXiv preprint arXiv:1606.08415, 2016.
    [7] C. Dong, C. C. Loy, K. He, and X. Tang, “Image super-resolution using deep convolutional networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 38, no. 2, pp. 295–307, 2016.
    [8] B. Lim, S. Son, H. Kim, S. Nah, and K. M. Lee, “Enhanced deep residual networks for single image super-resolution,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1132–1140, 2017.
    [9] W.-S. Lai, J.-B. Huang, N. Ahuja, and M.-H. Yang, “Deep laplacian pyramid networks for fast and accurate super-resolution,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5835–5843, 2017.
    [10] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations (ICLR), 2021.
    [11] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte, “SwinIR: Image restoration using swin transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), pp. 1833–1844, 2021.
    [12] L. Chen, X. Chu, X. Zhang, and J. Sun, “Simple baselines for image restoration,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 17–33, 2022.
    [13] H. Guo, J. Li, T. Dai, Z. Ouyang, X. Ren, and S.-T. Xia, “MambaIR: A simple baseline for image restoration with state-space model,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 222–241, 2024.
    [14] R. Li, R. T. Tan, and L.-F. Cheong, “All in one bad weather removal using architectural search,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3172–3182, 2020.
    [15] M. V. Conde, G. Geigle, and R. Timofte, “InstructIR: High-quality image restoration following human instructions,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 1–21, 2024.
    [16] B. Li, X. Liu, P. Hu, Z. Wu, J. Lv, and X. Peng, “All-in-one image restoration for unknown corruption,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 17 431–17 441, 2022.
    [17] J. Zhang, J. Huang, M. Yao, Z. Yang, H. Yu, M. Zhou, and F. Zhao, “Ingredient-oriented multi-degradation learning for image restoration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5825–5835, 2023.
    [18] E. Zamfir, Z. Wu, N. Mehta, Y. Tan, D. P. Paudel, Y. Zhang, and R. Timofte, “Complexity experts are task-discriminative learners for any image restoration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12 753–12 763, 2025.
    [19] Y. Cui, S. W. Zamir, M.-H. Yang, A. Knoll, F. S. Khan, and S. Khan, “StarIR: Convolutional image restoration with spatial-frequency fusion,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 48, no. 7, pp. 8216–8233, 2026.
    [20] Y. Cui, W. Ren, B. Shi, and A. Knoll, “Visual-in-visual: A unified and efficient baseline for image restoration,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 48, no. 7, pp. 7981–7999, 2026.
    [21] L. He, J. Chu, F. Lv, W. Liu, T. Li, J. Cheng, and Y. Fang, “Degradation-aware adaptive context gating for unified image restoration,” arXiv preprint arXiv:2605.01236, 2026.
    [22] X. Zhang, H. Zhang, G. Wang, Q. Zhang, and L. Zhang, “ClearAIR: A human-visual-perception-inspired all-in-one image restoration,” in AAAI Conference on Artificial Intelligence (AAAI), pp. 12 861–12 869, 2026.
    [23] S. Mallat, A Wavelet Tour of Signal Processing: The Sparse Way, 3rd ed. Academic Press, 2008.
    [24] J. Dong, J. Pan, Z. Yang, and J. Tang, “Multi-scale residual low-pass filter network for image deblurring,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 12 311–12 320, 2023.
    [25] O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), pp. 234–241, 2015.
    [26] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “LLaMA: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023.
    [27] T. P. Pires, A. V. Lopes, Y. Assogba, and H. Setiawan, “One wide feedforward is all you need,” in Proceedings of the Eighth Conference on Machine Translation (WMT), pp. 1031–1044, 2023.
    [28] X. Tang, X. Gu, X. He, X. Hu, and J. Sun, “Degradation-aware residual-conditioned optimal transport for unified image restoration,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 47, no. 8, pp. 6764–6779, 2025.
    [29] G. Wu, J. Jiang, K. Jiang, X. Liu, and L. Nie, “DSwinIR: Rethinking window-based attention for image restoration,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 48, pp. 4350–4366, 2026.
    [30] D. Serrano-Lozano, L. Herranz, S. Su, and J. Vazquez-Corral, “Adaptive blind all-in-one image restoration,” Computer Vision and Image Understanding (CVIU), vol. 269, p. 104795, 2026.
    [31] B. Ren, Y. Li, X. Zheng, Y. Fu, D. P. Paudel, H. Liu, M.-H. Yang, L. Van Gool, and N. Sebe, “Efficient degradation-agnostic image restoration via channel-wise functional decomposition and manifold regularization,” in International Conference on Learning Representations (ICLR), 2026.
    [32] H. Yang, H. Yin, M. Shen, P. Molchanov, H. Li, and J. Kautz, “Global vision transformer pruning with hessian-aware saliency,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18 547–18 557, 2023.
    [33] Y. Li, J. Hu, Y. Wen, G. Evangelidis, K. Salahi, Y. Wang, S. Tulyakov, and J. Ren, “Rethinking vision transformers for MobileNet size and speed,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 16 843–16 854, 2023.

    下載圖示
    校外:立即公開
    QR CODE