| 研究生: |
李明翰 Lee, Ming-Han |
|---|---|
| 論文名稱: |
nnMNet: 火星地形語意分割基準模型 nnMNet: Baseline for Martian Terrain Semantic Segmentation |
| 指導教授: |
陳奇業
Chen, Chi-Yeh |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 51 |
| 中文關鍵詞: | 火星 、深度學習 、語意分割 、地形測繪 |
| 外文關鍵詞: | Mars, deep learning, semantic segmentation, terrain mapping |
| 相關次數: | 點閱:7 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
語意分割在深入理解火星的過程中扮演著相當重要的角色。然而,由於火星表面是高度非結構化且極度複雜的,要實現精確的像素級預測與精細的資料標註相當困難。儘管近期的深度學習進展催生了許多方法與資料集以應對這些挑戰,該領域仍缺乏一個強大、公開且可重現的基準模型,以及一個統一且公平的基準測試。在本論文中,我們提出了一個專為火星地形語意分割設計的深度學習模型:nnMNet。基於 nnWNet 的架構,我們整合了線性注意力機制以更有效地捕捉全域上下文資訊,並採用輕量化卷積以降低運算開銷。為了彌補局部與全域特徵表示之間的差異,我們提出了一個高效的融合模組,用以增強並融合具備不同特性的特徵。此外,我們彙整並標準化三個高品質的資料集,建立全新的基準測試以進行全面性的性能評估。nnMNet 在 SynMars-TW、SynMars-Air 與 MarsScapes 資料集上,分別達到了86.61%、83.25% 與 88.24% 的 mIoU ,均為目前最佳的結果。程式碼、模型以及資料集將公開於:https://github.com/dereklee0310/nnMNet。
Semantic segmentation is a crucial task for understanding Mars, the most Earth-like planet in our solar system. However, it is challenging because the Martian surface is highly unstructured and complex, making accurate pixel-level prediction and fine-grained annotation difficult. Recent advancements in deep learning have introduced numerous methods and datasets to address these challenges. Nevertheless, the field lacks a robust, publicly available, and reproducible baseline, as well as a unified benchmark to facilitate fair evaluations. In this thesis, we present nnMNet, a new baseline model designed for Martian terrain semantic segmentation. Building upon nnWNet, we integrate linear attention to better capture global context and employ lightweight convolutions to reduce computational overhead. To bridge the gap between local and global representations, we introduce the Spatially-Aware Fusion Block (SAFB), which augments and combines features with diverse characteristics. Furthermore, we establish a new benchmark by curating and standardizing three high-quality datasets for thorough evaluation. nnMNet achieves new state-of-the-art 86.61%, 83.25%, and 88.24% mIoU on SynMars-TW, SynMars-Air, and MarsScapes, respectively. Our code, models, and datasets are publicly available at https://github.com/dereklee0310/nnMNet.
[1] Y. Qu, C.-O. Chow, J. H. Chuah, and K. L. Soon, “A systematic review of martian image segmentation techniques for mars exploration (from 2019 to 2025),” Advances in Space Research, 2025.
[2] D. Kass, J. Schofield, A. Kleinböhl, D. McCleese, N. Heavens, J. Shirley, and L. Steele, “Mars climate sounder observation of mars’ 2018 global dust storm,” Geophysical Research Letters, vol. 47, no. 23, e2019GL083931, 2020.
[3] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” International Conference on Learning Representations.
[4] H. Liu, M. Yao, X. Xiao, and Y. Xiong, “Rockformer: A u-shaped transformer network for martian rock segmentation,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–16, 2023.
[5] Y. Xiong, X. Xiao, M. Yao, H. Liu, H. Yang, and Y. Fu, “Marsformer: Martian rock semantic segmentation with transformer,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–12, 2023.
[6] Y. Xiong, X. Xiao, M. Yao, H. Cui, and Y. Fu, “Light4mars: A lightweight transformer model for semantic segmentation on unstructured environment like mars,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 214, pp. 167–178, 2024.
[7] Y. Qi, X. Xiao, M. Yao, Y. Xiong, L. Zhang, and H. Cui, “Airformer: Learning-based object detection for mars helicopter,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 100–111, 2024.
[8] B. Lin, F. Wang, Q. Li, B. Zheng, M. Yao, X. Xiao, Y. Qi, H. Cui, and X. Huang, “Lissemars: A lightweight semantic segmentation model for mars helicopter,” Aerospace, vol. 12, no. 12, p. 1049, 2025.
[9] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10 012–10 022.
[10] W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 568–578.
[11] W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pvt v2: Improved baselines with pyramid vision transformer,” Computational visual media, vol. 8, no. 3, pp. 415–424, 2022.
[12] E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “Segformer: Simple and efficient design for semantic segmentation with transformers,” Advances in neural information processing systems, vol. 34, pp. 12 077–12 090, 2021.
[13] Y. Zhou, L. Li, L. Lu, and M. Xu, “Nnwnet: Rethinking the use of transformers in biomedical image segmentation and calling for a unified evaluation benchmark,” Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 20 852–20 862.
[14] Q. Fan, H. Huang, and R. He, “Breaking the low-rank dilemma of linear attention,” Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 25 271–25 280.
[15] W. Luo, Y. Li, R. Urtasun, and R. Zemel, "Understanding the effective receptive field in deep convolutional neural networks," Advances in neural information processing systems, vol. 29, 2016.
[16] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, "Mobilenetv2: Inverted residuals and linear bottlenecks," Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4510–4520.
[17] F. Isensee, P. F. Jaeger, S. A. Kohl, J. Petersen, and K. H. Maier-Hein, "Nnu-net: A self-configuring method for deep learning-based biomedical image segmentation," Nature methods, vol. 18, no. 2, pp. 203–211, 2021.
[18] J. Long, E. Shelhamer, and T. Darrell, "Fully convolutional networks for semantic segmentation," Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3431–3440.
[19] O. Ronneberger, P. Fischer, and T. Brox, "U-net: Convolutional networks for biomedical image segmentation," International Conference on Medical image computing and computer-assisted intervention, Springer, 2015, pp. 234–241.
[20] V. Badrinarayanan, A. Kendall, and R. Cipolla, "Segnet: A deep convolutional encoder-decoder architecture for image segmentation," IEEE transactions on pattern analysis and machine intelligence, vol. 39, no. 12, pp. 2481–2495, 2017.
[21] H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, "Pyramid scene parsing network," Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2881–2890.
[22] L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, "Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 4, pp. 834–848, 2018.
[23] L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, "Encoder-decoder with atrous separable convolution for semantic image segmentation," Proceedings of the European conference on computer vision (ECCV), 2018, pp. 801–818.
[24] J. Fu, J. Liu, H. Tian, Y. Li, Y. Bao, Z. Fang, and H. Lu, "Dual attention network for scene segmentation," Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 3146–3154.
[25] M.-H. Guo, C.-Z. Lu, Q. Hou, Z. Liu, M.-M. Cheng, and S.-M. Hu, "Segnext: Rethinking convolutional attention design for semantic segmentation," Advances in neural information processing systems, vol. 35, pp. 1140–1156, 2022.
[26] B. Rothrock, R. Kennedy, C. Cunningham, J. Papon, M. Heverly, and M. Ono, "Spoc: Deep learning-based terrain classification for mars rover missions," AIAA space 2016, 2016, p. 5539.
[27] A. S. McEwen, E. M. Eliason, J. W. Bergstrom, N. T. Bridges, C. J. Hansen, W. A. Delamere, J. A. Grant, V. C. Gulick, K. E. Herkenhoff, L. Keszthelyi, et al., "Mars reconnaissance orbiter's high resolution imaging science experiment (hirise)," Journal of Geophysical Research: Planets, vol. 112, no. E5, 2007.
[28] J. P. Grotzinger, J. Crisp, A. R. Vasavada, R. C. Anderson, C. J. Baker, R. Barry, D. F. Blake, P. Conrad, K. S. Edgett, B. Ferdowski, et al., "Mars science laboratory mission and science investigation," Space science reviews, vol. 170, no. 1, pp. 5–56, 2012.
[29] K. He, G. Gkioxari, P. Dollár, and R. Girshick, "Mask r-cnn," Proceedings of the IEEE international conference on computer vision, 2017, pp. 2961–2969.
[30] K. Ogohara and R. Gichu, "Automated segmentation of textured dust storms on mars remote sensing images using an encoder-decoder type convolutional neural network," Computers & Geosciences, vol. 160, p. 105043, 2022.
[31] D. M. DeLatte, S. T. Crites, N. Guttenberg, E. J. Tasker, and T. Yairi, "Segmentation convolutional neural networks for automatic crater detection on mars," IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 12, no. 8, pp. 2944–2957, 2019.
[32] L. Rubanenko, S. Pérez-López, J. Schull, and M. G. Lapôtre, "Automatic detection and segmentation of barchan dunes on mars and earth using a convolutional neural network," IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 14, pp. 9364–9371, 2021.
[33] J. Hu, L. Shen, and G. Sun, "Squeeze-and-excitation networks," Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7132–7141.
[34] J. Park, S. Woo, J.-Y. Lee, and I.-S. Kweon, "Bam: Bottleneck attention module," British Machine Vision Conference, 2018.
[35] S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, "Cbam: Convolutional block attention module," Proceedings of the European conference on computer vision (ECCV), 2018, pp. 3–19.
[36] Q. Wang, B. Wu, P. Zhu, P. Li, W. Zuo, and Q. Hu, "Eca-net: Efficient channel attention for deep convolutional neural networks," Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 11 534–11 542.
[37] H. Liu, M. Yao, X. Xiao, and H. Cui, "A hybrid attention semantic segmentation network for unstructured terrain on mars," Acta Astronautica, vol. 204, pp. 492–499, 2023.
[38] D. Chen, F. Hu, P. T. Mathiopoulos, Z. Zhang, and J. Peethambaran, "Mc-unet: Martian crater segmentation at semantic and instance levels using u-net-based convolutional neural network," Remote Sensing, vol. 15, no. 1, p. 266, 2023.
[39] H. Li, L. Qiu, Z. Li, B. Meng, J. Huang, and Z. Zhang, "Automatic rocks segmentation based on deep learning for planetary rover images," Journal of Aerospace Information Systems, vol. 18, no. 11, pp. 755–761, 2021.
[40] L. Feng, S. Wang, D. Wang, P. Xiong, J. Xie, Y. Hu, M. Zhang, E. Q. Wu, and A. Song, "Mobile-deeprfb: A lightweight terrain classifier for automatic mars rover navigation," IEEE Transactions on Automation Science and Engineering, vol. 22, pp. 17 442–17 451, 2023.
[41] S. Liu, D. Huang, et al., "Receptive field block net for accurate and fast object detection," Proceedings of the European conference on computer vision (ECCV), 2018, pp. 385–400.
[42] A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan, W. Wang, Y. Zhu, R. Pang, V. Vasudevan, et al., "Searching for mobilenetv3," Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 1314–1324.
[43] J. Li, K. Chen, G. Tian, L. Li, and Z. Shi, "Marsseg: Mars surface semantic segmentation with multilevel extractor and connector," IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–12, 2025.
[44] V. Ashish, "Attention is all you need," Advances in neural information processing systems, vol. 30, p. I, 2017.
[45] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, "Training data-efficient image transformers & distillation through attention," International conference on machine learning, PMLR, 2021, pp. 10 347–10 357.
[46] Z. Dai, H. Liu, Q. V. Le, and M. Tan, "Coatnet: Marrying convolution and attention for all data sizes," Advances in neural information processing systems, vol. 34, pp. 3965–3977, 2021.
[47] Y. Dai, T. Zheng, C. Xue, and L. Zhou, "Segmarsvit: Lightweight mars terrain segmentation network for autonomous driving in planetary exploration," Remote Sensing, vol. 14, no. 24, p. 6297, 2022.
[48] Y. Jia, G. Wan, W. Li, C. Li, J. Liu, D. Cong, and L. Liu, "Edr-transunet: Integrating enhanced dual relation-attention with transformer u-net for multiscale rock segmentation on mars," IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–16, 2024.
[49] L. Fan, J. Yuan, X. Niu, K. Zha, and W. Ma, "Rockseg: A novel semantic segmentation network based on a hybrid framework combining a convolutional neural network and transformer for deep space rock images," Remote Sensing, vol. 15, no. 16, p. 3935, 2023.
[50] W. Lv, L. Wei, D. Zheng, Y. Liu, and Y. Wang, "Marsnet: Automated rock segmentation with transformers for tianwen-1 mission," IEEE Geoscience and Remote Sensing Letters, vol. 20, pp. 1–5, 2022.
[51] R. Wang, J. Sun, K. Zhou, J. Wang, J. Bi, Q. Zhang, W. Wang, G. Qu, C. Li, and H. Qiu, "Marsterrnet: A u-shaped dual-backbone framework with feature-guided loss for martian terrain segmentation," Remote Sensing, vol. 18, no. 1, p. 35, 2025.
[52] L. Fan, J. Yuan, and K. Zha, "Terseg: A dual-branch semantic segmentation network for mars terrain and autonomous path planning," Expert Systems with Applications, vol. 270, p. 126397, 2025.
[53] M. C. Malin, J. F. Bell III, B. A. Cantor, M. A. Caplinger, W. M. Calvin, R. T. Clancy, K. S. Edgett, L. Edwards, R. M. Haberle, P. B. James, et al., "Context camera investigation on board the mars reconnaissance orbiter," Journal of Geophysical Research: Planets, vol. 112, no. E5, 2007.
[54] M. Purohit, J. Adler, and H. Kerner, "Conequest: A benchmark for cone segmentation on mars," Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2024, pp. 6026–6035.
[55] S. Paheding, A. A. Reyes, A. Rajaneesh, K. Sajinkumar, and T. Oommen, "Marsls-net: Martian landslides segmentation network and benchmark dataset," Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 8236–8245.
[56] J. A. Crisp, M. Adler, J. R. Matijevic, S. W. Squyres, R. E. Arvidson, and D. M. Kass, "Mars exploration rover mission," Journal of Geophysical Research: Planets, vol. 108, no. E12, 2003.
[57] B. Wu, J. Dong, Y. Wang, W. Rao, Z. Sun, Z. Li, Z. Tan, Z. Chen, C. Wang, W. C. Liu, et al., "Landing site selection and characterization of tianwen-1 (zhurong rover) on mars," Journal of Geophysical Research: Planets, vol. 127, no. 4, e2021JE007137, 2022.
[58] X. Xiao, M. Yao, H. Liu, J. Wang, L. Zhang, and Y. Fu, "A kernel-based multi-featured rock modeling and detection framework for a mars rover," IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 7, pp. 3335–3344, 2021.
[59] X. Xiao, M. Yao, and H. Liu, "Marsdata-v2, a rock segmentation dataset of real martian scenes," IEEE Dataport. Available at: doi, vol. 10, 2022.
[60] R. M. Swan, D. Atha, H. A. Leopold, M. Gildner, S. Oij, C. Chiu, and M. Ono, "Ai4mars: A dataset for terrain-aware autonomous driving on mars," Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 1982–1991.
[61] J. Li, K. Chen, G. Tian, L. Li, and Z. Shi, "Marsseg: Mars surface semantic segmentation with multilevel extractor and connector," IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–12, 2025.
[62] J. Zhang, L. Lin, Z. Fan, W. Wang, and J. Liu, "S5mars: Semi-supervised learning for mars semantic segmentation," IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–15, 2024.
[63] H. Liu, M. Yao, X. Xiao, B. Zheng, and H. Cui, "Marsscapes and udaformer: A panorama dataset and a transformer-based unsupervised domain adaptation framework for martian terrain segmentation," IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–17, 2023.
[64] A. Bréhéret, Pixel Annotation Tool, 2017.
[65] J. Cartucho, S. Tukra, Y. Li, D. S. Elson, and S. Giannarou, "Visionblender: A tool to efficiently generate computer vision datasets for robotic surgery," Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization, pp. 1–8, 2020.
[66] T. Xiao, M. Singh, E. Mintun, T. Darrell, P. Dollár, and R. Girshick, "Early convolutions help transformers see better," Advances in neural information processing systems, vol. 34, pp. 30 392–30 400, 2021.
[67] B. Graham, A. El-Nouby, H. Touvron, P. Stock, A. Joulin, H. Jégou, and M. Douze, "Levit: A vision transformer in convnet's clothing for faster inference," Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 12 259–12 269.
[68] S. Mehta and M. Rastegari, "Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer," International Conference on Learning Representations.
[69] J. Guo, K. Han, H. Wu, Y. Tang, X. Chen, Y. Wang, and C. Xu, "Cmt: Convolutional neural networks meet vision transformers," Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 12 175–12 185.
[70] C. Yang, S. Qiao, Q. Yu, X. Yuan, Y. Zhu, A. Yuille, H. Adam, and L.-C. Chen, "Moat: Alternating mobile convolution and attention brings strong vision models," The Eleventh International Conference on Learning Representations.
[71] C. Si, W. Yu, P. Zhou, Y. Zhou, X. Wang, and S. Yan, "Inception transformer," Advances in neural information processing systems, vol. 35, pp. 23 495–23 509, 2022.
[72] Z. Peng, W. Huang, S. Gu, L. Xie, Y. Wang, J. Jiao, and Q. Ye, "Conformer: Local features coupling global representations for visual recognition," Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 367–376.
[73] Y. Chen, X. Dai, D. Chen, M. Liu, X. Dong, L. Yuan, and Z. Liu, "Mobile-former: Bridging mobilenet and transformer," Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 5270–5279.
[74] O. Oktay, J. Schlemper, L. Le Folgoc, M. Lee, M. Heinrich, K. Misawa, K. Mori, S. McDonagh, N. Y. Hammerla, B. Kainz, et al., "Attention u-net: Learning where to look for the pancreas," Medical Imaging with Deep Learning.
[75] Q. Fan, H. Huang, X. Zhou, and R. He, "Lightweight vision transformer with bidirectional interaction," Advances in Neural Information Processing Systems, vol. 36, pp. 15 234–15 251, 2023.
[76] W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, and S. Yan, "Metaformer is actually what you need for vision," Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10 819–10 829.
[77] K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
[78] S. Ioffe and C. Szegedy, "Batch normalization: Accelerating deep network training by reducing internal covariate shift," International conference on machine learning, pmlr, 2015, pp. 448–456.
[79] D. Hendrycks and K. Gimpel, "Gaussian error linear units (gelus)," arXiv preprint arXiv:1606.08415, 2016.
[80] D. Han, X. Pan, Y. Han, S. Song, and G. Huang, "Flatten transformer: Vision transformer using focused linear attention," Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 5961–5971.
[81] D. Han, Y. Pu, Z. Xia, Y. Han, X. Pan, X. Li, J. Lu, S. Song, and G. Huang, "Bridging the divide: Reconsidering softmax and linear attention," Advances in Neural Information Processing Systems, vol. 37, pp. 79 221–79 245, 2024.
[82] M. A. Islam, S. Jia, and N. D. Bruce, "How much position information do convolutional neural networks encode?" International Conference on Learning Representations.
[83] P. Shaw, J. Uszkoreit, and A. Vaswani, "Self-attention with relative position representations," Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), 2018, pp. 464–468.
[84] X. Chu, Z. Tian, B. Zhang, X. Wang, and C. Shen, "Conditional positional encodings for vision transformers," The Eleventh International Conference on Learning Representations.
[85] X. Dong, J. Bao, D. Chen, W. Zhang, N. Yu, L. Yuan, D. Chen, and B. Guo, "Cswin transformer: A general vision transformer backbone with cross-shaped windows," Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 12 124–12 134.
[86] K. Yuan, S. Guo, Z. Liu, A. Zhou, F. Yu, and W. Wu, "Incorporating convolution designs into visual transformers," Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 579–588.
[87] J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, "Roformer: Enhanced transformer with rotary position embedding," Neurocomputing, vol. 568, p. 127063, 2024.
[88] F. Wang, S. Ren, T. Zhang, P. Neskovic, A. Bhattad, C. Xie, and A. Yuille, "Vit-5: Vision transformers for the mid-2020s," arXiv preprint arXiv:2602.08071, 2026.
[89] I. Loshchilov and F. Hutter, "Decoupled weight decay regularization," International Conference on Learning Representations.
[90] M. Contributors, MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark, https://github.com/open-mmlab/mmsegmentation, 2020.
[91] C. Yu, C. Gao, J. Wang, G. Yu, C. Shen, and N. Sang, "Bisenet v2: Bilateral network with guided aggregation for real-time semantic segmentation," International journal of computer vision, vol. 129, no. 11, pp. 3051–3068, 2021.
[92] M.-H. Guo, C.-Z. Lu, Z.-N. Liu, M.-M. Cheng, and S.-M. Hu, "Visual attention network," Computational visual media, vol. 9, no. 4, pp. 733–752, 2023.
[93] X. Ding, X. Zhang, J. Han, and G. Ding, "Scaling up your kernels to 31x31: Revisiting large kernel design in cnns," Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 11 963–11 975.
[94] S. Liu, T. Chen, X. Chen, X. Chen, Q. Xiao, B. Wu, T. Kärkkäinen, M. Pechenizkiy, D. C. Mocanu, and Z. Wang, "More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity," The Eleventh International Conference on Learning Representations.
[95] K. W. Lau, L.-M. Po, and Y. A. U. Rehman, "Large separable kernel attention: Rethinking the large kernel attention design in cnn," Expert Systems with Applications, vol. 236, p. 121352, 2024.
[96] Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, "A convnet for the 2020s," Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 11 976–11 986.