| 研究生: |
黃詩哲 Huang, Shih-Je |
|---|---|
| 論文名稱: |
易於部署之線性時間複雜度查表式影像去霧網路 LDL-Net: An Edge-Friendly Linear-Time Lookup Network for Image Dehazing |
| 指導教授: |
林家祥
Lin, Chia-Hsiang |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電腦與通信工程研究所 Institute of Computer & Communication Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 61 |
| 中文關鍵詞: | 影像去霧 、三維查找表 、Mamba 、YUV色彩空間 、邊緣人工智慧 、深度學習 |
| 外文關鍵詞: | Image Dehazing, 3D Lookup Tables, Mamba, YUV Color Space, Edge Artificial Intelligence, Deep Learning |
| 相關次數: | 點閱:93 下載:1 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
近年來,戶外視覺系統廣泛應用於自動駕駛、智慧交通、遠端監測及影像監控等領域,並須在存在霧氣干擾的戶外環境下提供可靠的視覺資訊。然而,霧氣會造成光線散射與衰減,使影像產生對比降低、色彩失真及細節模糊等現象,不僅降低人眼辨識能力,也影響後續電腦視覺任務的準確性。因此,如何在兼顧影像品質與運算效率的前提下,有效恢復清晰影像,已成為影像去霧領域的重要研究課題。
為了解決上述問題,本研究提出一種基於查找表之線性時間複雜度影像去霧架構(LDL-Net),結合場景自適應三維查找表與特徵引導機制,以低計算成本完成色彩映射與霧化校正。首先,利用 Mamba 特徵編碼器擷取輸入影像之全域場景特徵,並預測多組基底查找表的自適應權重,建構符合當前場景之查找表以完成色彩轉換。接著,於 YUV 色彩空間中設計亮度校正模組,利用高斯濾波器與拉普拉斯濾波器的自適應融合,有效提升影像對比並保留細節資訊。最後,透過殘差式細節修復模組進一步恢復紋理與結構細節,以提升整體去霧品質與視覺自然度。
實驗結果顯示,所提出之 LDL-Net 於公開影像去霧資料集上,在峰值訊噪比及結構相似性等客觀評估指標均展現具競爭力的表現,同時兼具較低的模型參數量、運算複雜度與推論延遲。此外,在高解析度影像推論實驗中,LDL-Net 可提供超過 8 倍的推論加速效果,並隨影像解析度提升仍維持近似線性的運算時間成長,展現良好的可擴展性與即時處理能力。本研究亦於 NVIDIA Jetson Orin Nano 邊緣運算平台完成部署驗證,證實所提出方法在資源受限環境下仍能維持良好的去霧效果與推論效率,展現其於自動駕駛、智慧監控、無人載具及其他邊緣人工智慧視覺應用中的發展潛力。
Outdoor vision systems are increasingly applied to autonomous driving, intelligent transportation, remote sensing, and surveillance. Nevertheless, haze caused by atmospheric scattering and attenuation reduces image visibility, distorts colors, and obscures structural details. Conventional single-image dehazing methods based on handcrafted priors often lack adaptability to complex scenes, whereas modern deep neural networks generally require substantial computation and memory, limiting their use in real-time and resource-constrained Edge Artificial Intelligence (Edge AI) platforms.
To address these limitations, this thesis proposes a Linear Dehazing Lookup Network (LDL-Net), which combines a Mamba-based feature encoder, Adaptive Color Transformation (ACT), Luminance Correction (LC), and Detail Refinement (DR). The feature encoder extracts global scene representations with linear-time sequence modeling. Guided by these features, ACT adaptively fuses three basis three-dimensional lookup tables (3D LUTs) and applies trilinear interpolation to perform scene-dependent color transformation. LC then processes the luminance channel in the YUV color space by adaptively combining Laplacian and Gaussian kernels, thereby enhancing local contrast while suppressing noise and avoiding unnecessary chrominance changes. Finally, DR learns residual information to recover subtle textures and structural details.
Experiments on the RESIDE-6K dataset demonstrate that LDL-Net achieves a PSNR of 30.37 dB and an SSIM of 0.95 with only 0.68 million parameters, 16.48 GFLOPs, and an average inference time of 7.35 ms. In addition, LDL-Net exhibits approximately linear runtime growth with increasing image resolution and achieves more than an 8 times inference speedup for high-resolution image dehazing. Deployment on the NVIDIA Jetson Orin Nano achieves 55.6 ms per image, approximately 4082.81 MB of memory usage, and an average power consumption of 9.65 W. These results show that LDL-Net provides a favorable balance among restoration quality, computational efficiency, and practical Edge AI deployment capability.
[1] R. M. Haralick, K. Shanmugam, and I. H. Dinstein, “Textural features for image classification,” IEEE Transactions on Systems, Man, and Cybernetics, no. 6, pp. 610–621, 1973.
[2] G. Oliveira, X. Frazão, A. Pimentel, and B. Ribeiro, “Automatic graphic logo detection via fast region-based convolutional networks,” in 2016 International Joint Conference on Neural Networks (IJCNN). IEEE, 2016, pp. 985–991.
[3] Y. Zhang, A. Carballo, H. Yang, and K. Takeda, “Autonomous driving in adverse weather conditions: A survey,” CoRR, vol. abs/2112.08936, 2021. [Online]. Available: https://arxiv.org/abs/2112.08936
[4] G. Hu, Y. Yang, D. Yi, J. Kittler, W. Christmas, S. Z. Li, and T. Hospedales, “When face recognition meets with deep learning: An evaluation of convolutional neural networks for face recognition,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) Workshops, December 2015.
[5] P.-W. Tang, C.-H. Lin, and Y. Liu, “Transformer-driven inverse problem transform for fast blind hyperspectral image dehazing,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–14, Jan. 2024.
[6] C.-H. Lin and P.-W. Tang, “Inverse problem transform: Solving hyperspectral inpainting via deterministic compressed sensing,” in Proc. Workshop on Hyperspectral Imaging and Signal Processing: Evolution in Remote Sensing, Amsterdam, Netherlands, 24–26 Mar. 2021, pp. 1–5.
[7] K. He, J. Sun, and X. Tang, “Single image haze removal using dark channel prior,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 33, no. 12, pp. 2341–2353, 2010.
[8] X. Liu, Y. Ma, Z. Shi, and J. Chen, “Griddehazenet: Attention-based multi-scale network for image dehazing,” in 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 7313–7322.
[9] Y. Song, Z. He, H. Qian, and X. Du, “Vision transformers for single image dehazing,” IEEE Transactions on Image Processing, vol. 32, pp. 1927–1941, 2023.
[10] S. G. Narasimhan and S. K. Nayar, “Vision and the atmosphere,” International journal of computer vision, vol. 48, no. 3, pp. 233–254, 2002.
[11] E. Y. Lam, “Combining gray world and retinex theory for automatic white balance in digital photography,” in Proceedings of the Ninth International Symposium on Consumer Electronics, 2005.(ISCE 2005). IEEE, 2005, pp. 134–139.
[12] F. Xiao, J. Liu, Y. Huang, E. Cheng, and F. Yuan, “Neuromorphic computing network for underwater image enhancement and beyond,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–17, 2024.
[13] R. Szeliski, Computer vision: algorithms and applications. Springer Nature, 2022.
[14] M. R. Hajiaboli, M. O. Ahmad, and C. Wang, “An edge-adapting Laplacian kernel for nonlinear diffusion filters,” IEEE Transactions on Image Processing, vol. 21, no. 4, pp. 1561–1572, 2011.
[15] R. C. Gonzalez and R. E. Woods, Digital Image Processing. Prentice Hall, 2002.
[16] M. Podpora, G. P. Korbas, and A. Kawala-Janik, “YUV vs RGB-choosing a color space for human-machine interaction.” in Proc. Conference on Computer Science and Intelligence Systems (FedCSIS), Warsaw, Poland, Sep. 2014, pp. 29–34.
[17] B. G. Haskell, A. Puri, and A. N. Netravali, Digital Video. Springer, 2002.
[18] H. Zeng, J. Cai, L. Li, Z. Cao, and L. Zhang, “Learning image-adaptive 3D lookup tables for high performance photo enhancement in real-time,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 4, pp. 2058–2073, 2020.
[19] Z. Wang et al., “Image quality assessment: From error visibility to structural similarity,” IEEE Transactions on Image Processing, vol. 13, no. 4, pp. 600–612, 2004.
[20] I. Bakurov, M. Buzzelli, R. Schettini, M. Castelli, and L. Vanneschi, “Structural similarity index (SSIM) revisited: A data-driven approach,” Expert Systems with Applications, vol. 189, p. 116087, 2022.
[21] A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” arXiv preprint arXiv:2312.00752, 2023, [Online]. Available: https://arxiv.org/abs/2312.00752.
[22] Y. Liu, Y. Tian, Y. Zhao, H. Yu, L. Xie, Y. Wang, Q. Ye, J. Jiao, and Y. Liu, “Vmamba: Visual state space model,” Advances in neural information processing systems, vol. 37, pp. 103 031–103 063, 2024.
[23] A. Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations, 2021.
[24] K. Han et al., “A survey on vision transformer,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
[25] J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” arXiv preprint arXiv:1607.06450, 2016.
[26] C. Liu, H. Yang, J. Fu, and X. Qian, “4d lut: Learnable context-aware 4d lookup table for image enhancement,” IEEE Transactions on Image Processing, vol. 32, pp. 4742–4756, 2023.
[27] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778.
[28] D. Ulyanov, A. Vedaldi, and V. Lempitsky, “Deep image prior,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 Jun. 2018, pp. 9446–9454.
[29] C.-H. Lin, Y.-C. Lin, and P.-W. Tang, “ADMM-ADAM: A new inverse imaging framework blending the advantages of convex optimization and deep learning,” IEEE Transactions on Geoscience and Remote Sensinging, vol. 60, pp. 1–16, Sep. 2021.
[30] B. Li, W. Ren, D. Fu, D. Tao, D. Feng, W. Zeng, and Z. Wang, “Benchmarking single image dehazing and beyond,” in CVPR, 2018.
[31] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101, 2017, [Online]. Available: https://arxiv.org/abs/1711.05101.
[32] NVIDIA, “Jetson nano developer kit,” 2023, https://developer.nvidia.com/embedded/jetson-nano.
[33] J. Chen and X. Ran, “Deep learning with edge computing: A review,” Proceedings of the IEEE, vol. 107, no. 8, pp. 1655–1674, 2019.
[34] Z. Zhou, X. Chen, E. Li, L. Zeng, K. Luo, and J. Zhang, “Edge intelligence: Paving the last mile of artificial intelligence with edge computing,” Proceedings of the IEEE, vol. 107, no. 8, pp. 1738–1762, 2019.