| 研究生: |
姚泓逸 Yao, Hung-Yi |
|---|---|
| 論文名稱: |
一個應用於皮膚病灶分割的增強型超輕量化視覺Mamba U-Net An Enhanced Ultra-Lightweight Vision Mamba U-Net for Skin Lesion Segmentation |
| 指導教授: |
戴顯權
Tai, Shen-Chuan |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電腦與通信工程研究所 Institute of Computer & Communication Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 74 |
| 中文關鍵詞: | 皮膚病灶分割 、視覺Mamba 、狀態空間模型 、輕量化網路 、深度監督 |
| 外文關鍵詞: | Skin Lesion Segmentation, Vision Mamba, State Space Model, Lightweight Network, Deep Supervision |
| 相關次數: | 點閱:86 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
皮膚病灶分割是黑色素瘤電腦輔助診斷中的關鍵步驟,而在資源受限的平台上部署,網路須以極低成本維持精確度。UltraLight VM-UNet 結合 U 型架構與平行視覺 Mamba(Parallel Vision Mamba, PVM)層,達到 0.049M 參數量與 0.060 GFLOPs。
本論文提出一個增強型超輕量化視覺 Mamba U-Net,使凡是不需可學習參數即可達成的功能都不消耗參數:零參數轉換 PVM 層以零參數實現 ASP-VMUNet 的空洞掃描與輪轉位移,並以零填充或截斷取代線性投影;跨層級高效通道注意力橋接移除通道注意力中各階段的全連接層;輕量區塊重建淺層階段;訓練端則採用混合損失與動態衰減的深度監督。
所得網路僅需 0.016M 參數量與 0.034 GFLOPs,較基準模型分別減少 67.0% 與 43.3%,且於批次大小為一時速度為其 1.60×。其在 ISIC2017 與 ISIC2018 上的 Dice 係數與基準模型持平;在僅作為外部測試集的 PH² 上,敏感度為表列值中最高,重疊度則列第三。本論文的貢獻因而在於精確度相當前提下的資源節省。
Skin lesion segmentation is a key step in computer-aided melanoma diagnosis, and resource-constrained deployment demands accuracy at very low cost. UltraLight VM-UNet, a U-shaped network with Parallel Vision Mamba (PVM) layers, reaches 0.049M parameters and 0.060 GFLOPs.
This Thesis proposes an enhanced ultra-lightweight Vision Mamba U-Net in which functionality obtainable without learnable parameters uses none. A Zero-Parameter Transition PVM layer runs ASP-VMUNet's atrous scan and shift round at no parameter cost and replaces the linear projection with zero-padding or truncation; a Cross-Level Efficient Channel Attention bridge removes the channel attention's per-stage fully connected layers; lightweight blocks rebuild the shallow stages; and training adds a hybrid loss and dynamic deep supervision.
The network uses 0.016M parameters and 0.034 GFLOPs, 67.0% and 43.3% below the baseline, and runs 1.60× as fast at batch size one. It matches the baseline's Dice coefficient on ISIC2017 and ISIC2018; on the external PH² test set it attains the highest tabulated sensitivity and the third-best overlap. The contribution is resource saving at comparable accuracy.
[1] R. L. Siegel, A. N. Giaquinto, and A. Jemal, "Cancer statistics, 2024," CA: A Cancer Journal for Clinicians, vol. 74, no. 1, pp. 12–49, 2024.
[2] A. Esteva, B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, H. M. Blau, and S. Thrun, "Dermatologist-level classification of skin cancer with deep neural networks," Nature, vol. 542, no. 7639, pp. 115–118, 2017.
[3] O. Ronneberger, P. Fischer, and T. Brox, "U-Net: Convolutional networks for biomedical image segmentation," in Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), pp. 234–241, 2015.
[4] K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778, 2016.
[5] O. Oktay, J. Schlemper, L. Le Folgoc, M. Lee, M. Heinrich, K. Misawa, K. Mori, S. McDonagh, N. Y. Hammerla, B. Kainz, B. Glocker, and D. Rueckert, "Attention U-Net: Learning where to look for the pancreas," in Proceedings of the International Conference on Medical Imaging with Deep Learning (MIDL), 2018.
[6] Z. Zhou, M. M. R. Siddiquee, N. Tajbakhsh, and J. Liang, "UNet++: A nested U-Net architecture for medical image segmentation," in Proceedings of the International Workshop on Deep Learning in Medical Image Analysis (DLMIA, MICCAI), pp. 3–11, 2018.
[7] J. Chen, Y. Lu, Q. Yu, X. Luo, E. Adeli, Y. Wang, L. Lu, A. L. Yuille, and Y. Zhou, "TransUNet: Transformers make strong encoders for medical image segmentation," arXiv preprint arXiv:2102.04306, 2021.
[8] H. Cao, Y. Wang, J. Chen, D. Jiang, X. Zhang, Q. Tian, and M. Wang, "Swin-Unet: Unet-like pure transformer for medical image segmentation," in Proceedings of the European Conference on Computer Vision Workshops (ECCVW), pp. 205–218, 2022.
[9] A. Gu, K. Goel, and C. Ré, "Efficiently modeling long sequences with structured state spaces," in Proceedings of the International Conference on Learning Representations (ICLR), 2022.
[10] A. Gu and T. Dao, "Mamba: Linear-time sequence modeling with selective state spaces," in Proceedings of the Conference on Language Modeling (COLM), 2024.
[11] L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang, "Vision Mamba: Efficient visual representation learning with bidirectional state space model," in Proceedings of the International Conference on Machine Learning (ICML), vol. 235, pp. 62429–62442, 2024.
[12] Y. Liu, Y. Tian, Y. Zhao, H. Yu, L. Xie, Y. Wang, Q. Ye, J. Jiao, and Y. Liu, "VMamba: Visual state space model," in Advances in Neural Information Processing Systems (NeurIPS), vol. 37, pp. 103031–103063, 2024.
[13] J. Ma, F. Li, and B. Wang, "U-Mamba: Enhancing long-range dependency for biomedical image segmentation," arXiv preprint arXiv:2401.04722, 2024.
[14] J. Ruan, J. Li, and S. Xiang, "VM-UNet: Vision Mamba UNet for medical image segmentation," ACM Transactions on Multimedia Computing, Communications, and Applications, Art. no. 3767748, 2025.
[15] R. Wu, Y. Liu, G. Ning, P. Liang, and Q. Chang, "UltraLight VM-UNet: Parallel Vision Mamba significantly reduces parameters for skin lesion segmentation," Patterns, vol. 6, no. 11, Art. no. 101298, 2025.
[16] M. Bao, S. Lyu, Z. Xu, Q. Zhao, C. Zeng, W. Bai, and G. Cheng, "ASP-VMUNet: Atrous shifted parallel Vision Mamba U-Net for skin lesion segmentation," arXiv preprint arXiv:2503.19427, 2025.
[17] Q. Wang, B. Wu, P. Zhu, P. Li, W. Zuo, and Q. Hu, "ECA-Net: Efficient channel attention for deep convolutional neural networks," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11534–11542, 2020.
[18] N. Abraham and N. M. Khan, "A novel focal Tversky loss function with improved attention U-Net for lesion segmentation," in Proceedings of the IEEE International Symposium on Biomedical Imaging (ISBI), pp. 683–687, 2019.
[19] C.-Y. Lee, S. Xie, P. Gallagher, Z. Zhang, and Z. Tu, "Deeply-supervised nets," in Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS), vol. 38, pp. 562–570, 2015.
[20] J. Long, E. Shelhamer, and T. Darrell, "Fully convolutional networks for semantic segmentation," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3431–3440, 2015.
[21] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, "Attention is all you need," in Advances in Neural Information Processing Systems (NeurIPS), pp. 5998–6008, 2017.
[22] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, "An image is worth 16x16 words: Transformers for image recognition at scale," in Proceedings of the International Conference on Learning Representations (ICLR), 2021.
[23] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, "Swin Transformer: Hierarchical vision transformer using shifted windows," in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 10012–10022, 2021.
[24] F. Chollet, "Xception: Deep learning with depthwise separable convolutions," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1251–1258, 2017.
[25] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, "MobileNets: Efficient convolutional neural networks for mobile vision applications," arXiv preprint arXiv:1704.04861, 2017.
[26] F. Yu and V. Koltun, "Multi-scale context aggregation by dilated convolutions," in Proceedings of the International Conference on Learning Representations (ICLR), 2016.
[27] L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, "DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 4, pp. 834–848, 2018.
[28] C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, "Going deeper with convolutions," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1–9, 2015.
[29] J. Ruan, S. Xiang, M. Xie, T. Liu, and Y. Fu, "MALUNet: A multi-attention and light-weight UNet for skin lesion segmentation," in Proceedings of the IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pp. 1150–1156, 2022.
[30] J. Ruan, M. Xie, J. Gao, T. Liu, and Y. Fu, "EGE-UNet: An efficient group enhanced UNet for skin lesion segmentation," in Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), pp. 481–490, 2023.
[31] X. Pei, T. Huang, and C. Xu, "EfficientVMamba: Atrous selective scan for light weight visual Mamba," in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), vol. 39, no. 6, pp. 6443–6451, 2025.
[32] W. Liao, Y. Zhu, X. Wang, C. Pan, Y. Wang, and L. Ma, "LightM-UNet: Mamba assists in lightweight UNet for medical image segmentation," arXiv preprint arXiv:2403.05246, 2024.
[33] Q.-H. Ho, T.-N.-Q. Nguyen, T.-T. Tran, and V.-T. Pham, "LiteMamba-Bound: A lightweight Mamba-based model with boundary-aware and normalized active contour loss for skin lesion segmentation," Methods, vol. 235, pp. 10–25, 2025.
[34] X. Li, W. Wang, X. Hu, and J. Yang, "Selective kernel networks," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 510–519, 2019.
[35] J. Hu, L. Shen, and G. Sun, "Squeeze-and-excitation networks," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7132–7141, 2018.
[36] S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, "CBAM: Convolutional block attention module," in Proceedings of the European Conference on Computer Vision (ECCV), pp. 3–19, 2018.
[37] F. Milletari, N. Navab, and S.-A. Ahmadi, "V-Net: Fully convolutional neural networks for volumetric medical image segmentation," in Proceedings of the International Conference on 3D Vision (3DV), pp. 565–571, 2016.
[38] S. S. M. Salehi, D. Erdogmus, and A. Gholipour, "Tversky loss function for image segmentation using 3D fully convolutional deep networks," in Proceedings of the International Workshop on Machine Learning in Medical Imaging (MLMI, MICCAI), pp. 379–387, 2017.
[39] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, "Focal loss for dense object detection," in Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 2980–2988, 2017.
[40] N. C. F. Codella, D. Gutman, M. E. Celebi, B. Helba, M. A. Marchetti, S. W. Dusza, A. Kalloo, K. Liopyris, N. Mishra, H. Kittler, and A. Halpern, "Skin lesion analysis toward melanoma detection: A challenge at the 2017 International Symposium on Biomedical Imaging (ISBI), hosted by the International Skin Imaging Collaboration (ISIC)," in Proceedings of the IEEE International Symposium on Biomedical Imaging (ISBI), pp. 168–172, 2018.
[41] N. Codella, V. Rotemberg, P. Tschandl, M. E. Celebi, S. Dusza, D. Gutman, B. Helba, A. Kalloo, K. Liopyris, M. Marchetti, H. Kittler, and A. Halpern, "Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the International Skin Imaging Collaboration (ISIC)," arXiv preprint arXiv:1902.03368, 2019.
[42] P. Tschandl, C. Rosendahl, and H. Kittler, "The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions," Scientific Data, vol. 5, Art. no. 180161, 2018.
[43] T. Mendonça, P. M. Ferreira, J. S. Marques, A. R. S. Marçal, and J. Rozeira, "PH² – A dermoscopic image database for research and benchmarking," in Proceedings of the Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), pp. 5437–5440, 2013.
[44] I. Loshchilov and F. Hutter, "Decoupled weight decay regularization," in Proceedings of the International Conference on Learning Representations (ICLR), 2019.
[45] I. Loshchilov and F. Hutter, "SGDR: Stochastic gradient descent with warm restarts," in Proceedings of the International Conference on Learning Representations (ICLR), 2017.