| 研究生: |
李家銘 Lee, Chia-Ming |
|---|---|
| 論文名稱: |
深度摺紙網路:摺疊與展開影像超解析度中的自相似性 Deep Origami Network: Folding and Unfolding Self-Similarity for Image Super-resolution |
| 指導教授: |
許志仲
Hsu, Chih-Chung 鄭順林 Jeng, Shuen-Lin |
| 學位類別: |
碩士 Master |
| 系所名稱: |
管理學院 - 數據科學研究所 Institute of Data Science |
| 論文出版年: | 2025 |
| 畢業學年度: | 113 |
| 語文別: | 中文 |
| 論文頁數: | 65 |
| 中文關鍵詞: | 深度摺紙網路 、單圖像超解析度 、自相似性 、幾何先驗 |
| 外文關鍵詞: | Deep Origami Network, Single Image Super-resolution, Self-similarity, Geometric Prior |
| 相關次數: | 點閱:116 下載:1 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
利用圖像的自相似性是單圖像超解析度(SISR)的關鍵,然而當前主流架構存在根本性限制:卷積神經網路(CNNs)受限於局部感受野,而 Vision Transformers 與 Mambas 等全域模型則因其一維序列扁平化的處理方式,破壞了捕捉幾何對稱性所必需的二維空間結構。我們為此提出深度摺紙網路(DON),一個嵌入了二面體群 $D_4$ 幾何先驗的新範式。其核心為一個模型無關的摺紙網路(ON)模組,它透過模擬摺紙操作,遞迴地聚合遠距離的對稱圖像塊。當 ON 與任何全域骨幹網路整合時,能賦予模型強大的二維幾何感知能力,從而能夠捕捉廣義的自相似性。我們進一步建立了層級式主資訊(PIR)框架,此框架揭示了一個雙層級優化原則(DOP):在輸入層級探索可重建的資訊(追求高 I-PIR),同時在學習到的特徵中最小化冗餘(追求低 F-PIR)。實驗證明,DON 模型在達到頂尖性能的同時,展現出極低的 F-PIR 值,此結果驗證了我們的理論,並實現了高效的特徵解耦。
Exploiting image self-similarity is key to Single Image Super-Resolution (SISR), yet prevailing architectures are fundamentally limited: CNNs by local receptive fields, and global models like Vision Transformers by their 1D sequence flattening, which disrupts the 2D spatial structure essential for capturing geometric symmetries. We introduce the Deep Origami Network (DON), a paradigm that embeds geometric priors from the dihedral group $D_4$. At its core, the model-agnostic Origami Network (ON) module recursively aggregates distant symmetric patches via simulated origami operations. When integrated with any global backbone, ON confers potent 2D geometric awareness, enabling the capture of generalized self-similarities. We further establish the Hierarchical Prime Information (PIR) framework, which reveals a Dual-level Optimization Principle (DOP) that explores reconstructible information at the input level (high I-PIR) while minimizing redundancy in the learned features (low F-PIR). Experiments demonstrate that DON models achieve state-of-the-art performance while exhibiting remarkably low F-PIR, validating our theory and achieving efficient feature decoupling.
[1] Eirikur Agustsson and Radu Timofte. Ntire 2017 challenge on single image super-resolution: Dataset and study. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, July 2017. 1
[2] T Bedford, F.M Dekking, M Breeuwer, M.S Keane, and D van Schooneveld. Fractal coding of monochrome images. Signal Processing: Image Communication, 6(5):405-419, 1994. 2
[3] Marco Bevilacqua, Aline Roumy, Christine Guillemot, and Marie Line Alberi-Morel. Low-complexity single-image super-resolution based on nonnegative neighbor embedding. 2012. 3
[4] Peng et al. Bo. RWKV: Reinventing RNNs for the transformer era. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023. 4
[5] Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. Pre-trained image processing transformer, 2021. 5
[6] Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(4):834-848, 2018. 6
[7] Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao, and Chao Dong. Activating more pixels in image super-resolution transformer, 2023. 7
[8] Zheng Chen, Yulun Zhang, Jinjin Gu, Yongbing Zhang, Linghe Kong, and Xin Yuan. Cross aggregation transformer for image restoration. In NeurIPS, 2022. 8
[9] Taco Cohen and Max Welling. Group equivariant convolutional networks. In Maria Florina Balcan and Kilian Q. Weinberger, editors, Proceedings of The 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, pages 2990-2999, New York, New York, USA, 20-22 Jun 2016. PMLR. 9
[10] Tri Dao and Albert Gu. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. In International Conference on Machine Learning (ICML), 2024. 10
[11] Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. Image super-resolution using deep convolutional networks, 2015. 11
[12] Alexey Dosovitskiy et al. An image is worth 16x16 words: Transformers for image recognition at scale, 2021. 12
[13] Yusuke Matsui et al. Sketch-based manga retrieval using manga109 dataset. Multimedia Tools and Applications, 76(20):21811-21838, 2017. 13
[14] Daniel Glasner, Shai Bagon, and Michal Irani. Super-resolution from a single image. In 2009 IEEE 12th International Conference on Computer Vision, pages 349-356, 2009. 14
[15] Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023. 15
[16] Jinjin Gu and Chao Dong. Interpreting super-resolution networks with local attribution maps. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9199-9208, 2021. 16
[17] Hang Guo, Yong Guo, Yaohua Zha, Yulun Zhang, Wenbo Li, Tao Dai, Shu-Tao Xia, and Yawei Li. Mambairv2: Attentive state space restoration. arXiv preprint arXiv:2411.15269, 2024. 17
[18] Hang Guo, Jinmin Li, Tao Dai, Zhihao Ouyang, Xudong Ren, and Shu-Tao Xia. Mambair: A simple baseline for image restoration with state-space model. In ECCV, 2024. 18
[19] Lingshen He, Yuxuan Chen, Zhengyang Shen, Yiming Dong, Yisen Wang, and Zhouchen Lin. Efficient equivariant network. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021. 19
[20] Chih-Chung Hsu, Chia-Ming Lee, and Yi-Shiuan Chou. Drct: Saving image super-resolution away from information bottleneck. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 6133-6142, June 2024. 20
[21] Jia-Bin Huang, Abhishek Singh, and Narendra Ahuja. Single image super-resolution from transformed self-exemplars. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5197-5206, 2015. 21
[22] Jia-Bin Huang, Abhishek Singh, and Narendra Ahuja. Single image super-resolution from transformed self-exemplars. In CVPR, 2015. 22
[23] Jiwon Kim, Jung Kwon Lee, and Kyoung Mu Lee. Accurate image super-resolution using very deep convolutional networks. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1646-1654, 2016. 23
[24] Wenbo Li, Xin Lu, Shengju Qian, Jiangbo Lu, Xiangyu Zhang, and Jiaya Jia. On efficient transformer and image pre-training for low-level vision. arXiv preprint arXiv:2112.10175, 2021. 24
[25] Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. Swinir: Image restoration using swin transformer. arXiv preprint arXiv:2108.10257, 2021. 25
[26] Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. Enhanced deep residual networks for single image super-resolution. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, July 2017. 26
[27] Wenjie Luo, Yujia Li, Raquel Urtasun, and Richard Zemel. Understanding the effective receptive field in deep convolutional neural networks, 2017. 27
[28] Benoit Mandelbrot. How long is the coast of britain? statistical self-similarity and fractional dimension. Science, 156(3775):636-638, May 1967. 28
[29] Benoit Mandelbrot. The Fractal Geometry of Nature. San Francisco, CA, 1982. 29
[30] D. Martin, C. Fowlkes, D. Tal, and J. Malik. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In Proceedings Eighth IEEE International Conference on Computer Vision. ICCV 2001, volume 2, pages 416-423 vol.2, 2001. 30
[31] Yiqun Mei, Yuchen Fan, Yuqian Zhou, Lichao Huang, Thomas S Huang, and Humphrey Shi. Image super-resolution with cross-scale non-local attention and exhaustive self-exemplars mining. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 31
[32] Guotao Meng, Yue Wu, and Qifeng Chen. Improving video super-resolution with long-term self-exemplars. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 5992-5998, 2023. 32
[33] Michael Poli, Winnie Xu, Stefano Massaroli, Chenlin Meng, Kuno Kim, and Stefano Ermon. Self-similarity priors: Neural collages as differentiable fractal representations. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022. 33
[34] Eduardo Pérez-Pellitero, Jordi Salvador, Javier Ruiz-Hidalgo, and Bodo Rosenhahn. Psyco: Manifold span reduction for super resolution. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1837-1845, 2016. 34
[35] Qing Qu, Ju Sun, and John Wright. Finding a sparse vector in a subspace: Linear sparsity using alternating directions. IEEE Transactions on Information Theory, 62(10):5855-5880, October 2016. 35
[36] Yuchuan Tian, Hanting Chen, Chao Xu, and Yunhe Wang. Image processing gnn: Breaking rigidity in super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 24108-24117, June 2024. 36
[37] Radu Timofte, Eirikur Agustsson, Luc Van Gool, Ming-Hsuan Yang, Lei Zhang, Bee Lim, et al. Ntire 2017 challenge on single image super-resolution: Methods and results. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, July 2017. 37
[38] Chien-Yao Wang, Hong-Yuan Mark Liao, and I-Hau Yeh. Designing network design strategies through gradient path analysis. arXiv preprint arXiv:2211.04800, 2022. 38
[39] Jianchao Yang, John Wright, Thomas S. Huang, and Yi Ma. Image super-resolution via sparse representation. IEEE Transactions on Image Processing, 19(11):2861-2873, 2010. 39
[40] Maoke Yang, Kun Yu, Chi Zhang, Zhiwei Li, and Kuiyuan Yang. Denseaspp for semantic segmentation in street scenes. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3684-3692, 2018. 40
[41] Du Yuzhen, Hu Teng, Zhang Jiangning, Yi Ran, Xu Chengming, Hu Xiaobin, Wu Kai, Luo Donghao, Wang Yabiao, and Ma Lizhuang. Exploring realsynthetic dataset and linear attention in image restoration, 2024. 41
[42] Roman Zeyde, Michael Elad, and Matan Protter. On single image scale-up using sparse-representations. In International conference on curves and surfaces, pages 711-730. Springer, 2010. 42
[43] Kaibing Zhang, Xinbo Gao, Dacheng Tao, and Xuelong Li. Multi-scale dictionary for single image super-resolution. In 2012 IEEE Conference on Computer Vision and Pattern Recognition, pages 1114-1121, 2012. 43
[44] Leheng Zhang, Yawei Li, Xingyu Zhou, Xiaorui Zhao, and Shuhang Gu. Transcending the limit of local window: Advanced super-resolution transformer with adaptive token dictionary. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2856-2865, June 2024. 44
[45] Yulun Zhang, Kunpeng Li, Kai Li, Lichen Wang, Bineng Zhong, and Yun Fu. Image super-resolution using very deep residual channel attention networks. In ECCV, 2018. 45
[46] Qiang Zhu, Pengfei Li, and Qianhui Li. Attention retractable frequency fusion transformer for image super resolution. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 1756-1763, 2023. 46