| 研究生: |
楊丞勛 Yang, Cheng-Syun |
|---|---|
| 論文名稱: |
一個輕量事件導引融合之卷積動態去模糊法 A Lightweight Event-Guided Fusion Approach for Convolutional Motion Deblurring |
| 指導教授: |
戴顯權
Tai, Shen-Chuan |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 64 |
| 中文關鍵詞: | 動態去模糊 、事件相機 、跨模態融合 、輕量化網路 |
| 外文關鍵詞: | motion deblurring, event camera, cross-modal fusion, lightweight network |
| 相關次數: | 點閱:76 下載:1 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
動態模糊是影像復原領域中常見的退化問題,由曝光期間相機或物體的相對運動所造成,會導致邊緣模糊與細節喪失,影響後續電腦視覺系統的準確性。現有的單影像去模糊方法僅依賴模糊影像本身進行復原,在面對嚴重運動模糊時往往難以準確還原清晰細節。事件相機能以微秒級的時間解析度非同步捕捉亮度變化,為去模糊任務提供豐富的互補運動資訊。然而,現有的事件導引去模糊方法通常需要從零設計專用的跨模態架構,增加了模型複雜度與設計成本。
為此,本論文提出一種輕量化的事件導引融合方法,以極少的額外參數將事件資訊整合至現有的卷積影像復原網路中。所提方法設計了多尺度事件編碼器與事件導引融合模組兩個核心元件。前者從事件表示中提取多尺度特徵,使其與網路各層的影像特徵自然對齊;後者透過通道門控機制自適應地控制事件資訊的注入量,在需要引導的區域有效利用運動資訊,在靜止區域則自動忽略。整個融合模組僅佔原始網路參數量不到百分之二點五,無需重新設計完整架構即可顯著提升去模糊效能。
實驗結果顯示,所提方法在合成與真實事件資料集上的整體表現優於原始架構與多個具代表性的現有方法,顯示其具備良好的實用性與有效性。
Motion blur is a common degradation in image restoration, caused by relative motion during exposure, leading to edge distortion and detail loss that affect downstream computer vision systems. Existing single image deblurring methods rely solely on the blurred image, which often contains insufficient information for accurate recovery under severe motion blur. Event cameras asynchronously capture brightness changes with microsecond-level temporal resolution, providing rich complementary motion information. However, available event-guided deblurring methods typically require dedicated cross-modal architectures designed from scratch, increasing model complexity and design cost.
To address this, this Thesis proposes a lightweight event-guided fusion approach that integrates event camera information into an existing convolutional image restoration network with minimal additional parameters. The proposed method introduces two core components: a Multi-Scale Event Encoder that extracts hierarchical features from the event representation aligned with each spatial scale of the network, and an Event-Guided Fusion module that employs a channel-wise gating mechanism to adaptively control event information injection, effectively utilizing motion cues where beneficial while suppressing them in static regions. The entire fusion module accounts for less than 2.5% of the original network parameters, significantly improving deblurring performance without redesigning the complete architecture.
Experimental results demonstrate that the proposed method performs favorably against the original architecture and several representative existing methods on both synthetic and real-world datasets, suggesting its effectiveness and practicality.
[1] S. Nah, T. Hyun Kim, and K. Mu Lee, “Deep multi-scale convolutional neural network for dynamic scene deblurring,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3883–3891, 2017.
[2] S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, M.-H. Yang, and L. Shao, “Multi-stage progressive image restoration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 14 821–14 831, 2021.
[3] S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M.-H. Yang, “Restormer: Efficient transformer for high-resolution image restoration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5728–5739, 2022.
[4] Y. Cui, W. Ren, X. Cao, and A. Knoll, “Revitalizing convolutional network for image restoration,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 9423–9438, 2024.
[5] L. Sun, C. Sakaridis, J. Liang, Q. Jiang, K. Yang, P. Sun, Y. Ye, K. Wang, and L. Van Gool, “Event-based fusion for motion deblurring with cross-modal attention,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 412–428,2022.
[6] Q. Chu, Y. Wang, Y. Zhang, and Y. Jiang, “Event-based image deblurring via cross-modal interaction fusion,” IEEE Transactions on Industrial Informatics, vol. 22, no. 2, pp. 1292–1301, 2026.
[7] O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), pp. 234–241, 2015.
[8] R. Fergus, B. Singh, A. Hertzmann, S. T. Roweis, and W. T. Freeman, “Removing camera shake from a single photograph,” in ACM SIGGRAPH 2006 Papers, pp. 787–794, 2006.
[9] S. Cho and S. Lee, “Fast motion deblurring,” ACM Transactions on Graphics (TOG), vol. 28, no. 5, pp. 1–8, 2009.
[10] Q. Shan, J. Jia, and A. Agarwala, “High-quality motion deblurring from a single image,” ACM Transactions on Graphics (TOG), vol. 27, no. 3, pp. 1–10, 2008.
[11] J. Pan, D. Sun, H. Pfister, and M.-H. Yang, “Blind image deblurring using dark channel prior,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1628–1636, 2016.
[12] O. Whyte, J. Sivic, A. Zisserman, and J. Ponce, “Non-uniform deblurring for shaken images,” International Journal of Computer Vision, vol. 98, no. 2, pp. 168–186, 2012.
[13] J. Sun, W. Cao, Z. Xu, and J. Ponce, “Learning a convolutional neural network for non-uniform motion blur removal,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 769–777, 2015.
[14] X. Tao, H. Gao, X. Shen, J. Wang, and J. Jia, “Scale-recurrent network for deep image deblurring,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8174–8182, 2018.
[15] O. Kupyn, V. Budzan, M. Mykhailych, D. Mishkin, and J. Matas, “DeblurGAN: Blind motion deblurring using conditional adversarial networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8183–8192, 2018.
[16] S.-J. Cho, S.-W. Ji, J.-P. Hong, S.-W. Jung, and S.-J. Ko, “Rethinking coarse-to-fine approach in single image deblurring,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 4641–4650, 2021.
[17] F.-J. Tsai, Y.-T. Peng, Y.-Y. Lin, C.-C. Tsai, and C.-W. Lin, “Stripformer: Strip trans-former for fast image deblurring,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 146–162, 2022.
[18] L. Kong, J. Dong, J. Ge, M. Li, and J. Pan, “Efficient frequency domain-based transformers for high-quality image deblurring,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5886–5895, 2023.
[19] L. Chen, X. Chu, X. Zhang, and J. Sun, “Simple baselines for image restoration,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 17–33, 2022.
[20] X. Jiang, N. Gao, H. Dou, X. Zhang, X. Zhong, Y. Deng, and H. Li, “Global modeling matters: A fast, lightweight, and effective baseline for efficient image restoration,” IEEE Transactions on Image Processing, vol. 35, pp. 2740–2754, 2026.
[21] L. Kong, J. Dong, J. Tang, M.-H. Yang, and J. Pan, “Efficient visual state space model for image deblurring,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12 710–12 719, 2025.
[22] G. Gallego, T. Delbruck, G. Orchard, C. Bartolozzi, B. Taba, A. Censi, S. Leutenegger, A. J. Davison, J. Conradt, K. Daniilidis et al.,“Event-based vision: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 1, pp. 154–180, 2020.
[23] P. Lichtsteiner, C. Posch, and T. Delbruck, “A 128× 128 120 dB 15 𝜈s latency asynchronous temporal contrast vision sensor,” IEEE Journal of Solid-State Circuits, vol. 43, no. 2, pp. 566–576, 2008.
[24] C. Brandli, R. Berner, M. Yang, S.-C. Liu, and T. Delbruck, “A 240×180 130 dB 3 𝜈s latency global shutter spatiotemporal vision sensor,” IEEE Journal of Solid-State Circuits, vol. 49, no. 10, pp. 2333–2341, 2014.
[25] X. Lagorce, G. Orchard, F. Galluppi, B. E. Shi, and R. B. Benosman, “HOTS: A hierarchy of event-based time-surfaces for pattern recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 7, pp. 1346–1359, 2017.
[26] A. Z. Zhu, L. Yuan, K. Chaney, and K. Daniilidis, “Unsupervised event-based learning of optical flow, depth, and egomotion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 989–997, 2019.
[27] D. Gehrig, A. Loquercio, K. G. Derpanis, and D. Scaramuzza, “End-to-end learning of representations for asynchronous event-based data,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 5632–5642, 2019.
[28] L. Pan, C. Scheerlinck, X. Yu, R. Hartley, M. Liu, and Y. Dai, “Bringing a blurry frame alive at high frame-rate with an event camera,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6820–6829, 2019.
[29] Z. Jiang, Y. Zhang, D. Zou, J. Ren, J. Lv, and Y. Liu, “Learning event-based motion deblurring,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3320–3329, 2020.
[30] W. Shang, D. Ren, D. Zou, J. S. Ren, P. Luo, and W. Zuo, “Bringing events into video deblurring with non-consecutively blurry frames,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 4531–4540, 2021.
[31] H. Rebecq, D. Gehrig, and D. Scaramuzza, “ESIM: An open event camera simulator,” in Proceedings of the Conference on Robot Learning (CoRL). PMLR, pp. 969–982, 2018.
[32] L. Chen, X. Lu, J. Zhang, X. Chu, and C. Chen, “HINet: Half instance normalization network for image restoration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 182–192, 2021.
[33] D. Hendrycks and K. Gimpel, “Gaussian error linear units (GELUs),” arXiv preprint arXiv:1606.08415,2016.
[34] J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7132–7141, 2018.
[35] H. Zhao, O. Gallo, I. Frosio, and J. Kautz, “Loss functions for image restoration with neural networks,” IEEE Transactions on Computational Imaging, vol. 3, no. 1, pp. 47–57, 2017.
[36] T. Stoffregen, C. Scheerlinck, D. Scaramuzza, T. Drummond, N. Barnes, L. Kleeman, and R. Mahony, “Reducing the sim-to-real gap for event cameras,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 534–549, 2020.
[37] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proceedings of the International Conference on Learning Representations (ICLR), 2015.
[38] I. Loshchilov and F. Hutter, “SGDR: Stochastic gradient descent with warm restarts,” in Proceedings of the International Conference on Learning Representations (ICLR), 2017.
[39] A. Tarvainen and H. Valpola, “Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results,” in Advances in Neural Information Processing Systems (NeurIPS), 2017.
[40] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE Transactions on Image Processing, vol. 13, no. 4, pp. 600–612, 2004.