簡易檢索 / 詳目顯示

研究生: 連思涵
LIEN, Szu-Han
論文名稱: 整合光學顯微鏡之新穎跨模態深度學習方法於二維材料的碎片分割與分類
A Novel Cross-Modal Deep Learning Approach Integrated with Optical Microscopy for Improved Segmentation and Fragment Classification in Two-Dimensional Materials
指導教授: 陳奇業
Chen, Chi-Yeh
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2025
畢業學年度: 113
語文別: 英文
論文頁數: 54
中文關鍵詞: 二維材料辨識 、基於自然語言的語意分割 、二元分類
外文關鍵詞: Two-dimensional Materials, Referring Expression Segmentation, Binary Classification
相關次數: 點閱:143  下載:1 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 對於二維材料的探索及其性質的研究,在近年來迅速發展,並廣泛應用於生物製藥、能源儲存、積體電路、行動通訊等領域。在使用設備製備高品質的二維材料時,材料的厚度與層數,對品質具有關鍵性的影響,因此產出合適的薄膜十分重要,目前常見的製備方法有兩種,分別為化學氣相沉積(CVD)以及機械剝離法,前者需要高溫環境,且涉及複雜的化學反應而易產生污染,後者高度依賴人工剝離技術,無法保證穩定性;此外,當我們獲得薄膜後,要判斷該材料面積是否夠大、夠完整以應用於後續製程,目前仍然是依賴專家主觀經驗進行判斷,缺乏自動化與客觀性的判定標準。

    本研究旨在運用深度學習的影像分割技術來協助辨識材料的單層,並進一步分類其對於製程的可行性。我們的研究引入了 Referring Expression Segmentation(RES)的方法來提升模型的理解能力,透過雙向跨模態融合模組產生融合特徵讓模型可以同時理解文字語意與視覺特徵,進而辨識出材料的性質,且在自然語言表達的輔助下,模型的分割準確率也進一步提升,讓分割出來的材料邊界更加清晰。我們也提出了 Successful Label 分類機制,藉此減少人工檢查的需求,增加自動化效率,同時使模型具備處理多任務的能力。為了使現實環境中的問題被解決,我們也將訓練後的模型整合至成大光電系的實驗室中,利用光學顯微鏡的影像達成即時辨識。

    The exploration of two-dimensional materials and their unique physical properties has progressed rapidly in recent years, with broad applications such as biopharmaceuticals, energy storage, integrated circuits, and mobile communications. During the fabrication of high-quality 2D materials using experimental equipment, the thickness and number of layers play a critical role in determining the overall quality of the materials. Therefore, producing sufficiently thin and uniform films is essential.

    Among the commonly used fabrication methods, Chemical Vapor Deposition (CVD) and mechanical exfoliation are the two most prevalent. While CVD requires high-temperature conditions and involves complex chemical reactions that are prone to contamination, mechanical exfoliation heavily relies on manual peeling techniques, leading to inconsistency and limited reproducibility. Moreover, even after obtaining exfoliated films, assessing whether a monolayer is sufficiently large and complete for downstream processing remains a subjective task, largely dependent on expert judgment, and currently lacks automated evaluation standards.

    In this thesis, our aim is to employ deep learning-based image segmentation techniques to facilitate the identification of monolayer regions in 2D materials and further determine their applicability in practical manufacturing processes. Specifically, we leverage Referring Expression Segmentation (RES) approach to enhance the model’s comprehension ability by integrating both visual and textual information. A bidirectional cross-modal fusion module is designed to generate fused features that jointly capture semantic and spatial cues, allowing the model to understand the characteristics of materials more effectively. With the assistance of natural language expressions, segmentation accuracy is significantly improved, resulting in clearer fragment boundaries.

    In addition, we propose a Successful Label classification mechanism to automate the assessment of a fragment’s usability, thus reducing the need for manual inspection and improving the overall efficiency of the workflow. This allows our model to handle multiple tasks, including segmentation and classification. Most importantly, to address practical challenges in real-world scenarios, we further integrate the trained model into the Photonics laboratory at National Cheng Kung University, enabling real-time recognition through optical microscopy systems.

    中文摘要 i Abstract ii 誌謝 iv Contents v List of Tables vii List of Figures viii 1 Introduction 1 1.1 Overview 1 1.2 Motivation 2 1.3 Contribution 4 2 Related Works 6 2.1 Advancements in Deep Learning for Two-Dimensional Materials 6 2.1.1 Image Classification Techniques 7 2.1.2 Image Segmentation Techniques 7 2.2 Referring Expression Segmentation 8 2.3 Binary Image Classification 9 3 Method 11 3.1 Encoder 12 3.1.1 Cross-Modal Fusion 12 3.1.2 Text Enhancement 15 3.2 Decoder 16 3.2.1 Segmentation 16 3.2.2 Classification Head 17 3.3 Loss Function 18 3.4 Integration of Deep Learning Models into Microscopy Systems 19 4 Performance Evaluation 21 4.1 Experimental Settings 21 4.1.1 Dataset 21 4.1.2 Evaluation Metric 22 4.1.3 Implementation Details 24 4.2 Experimental Results 25 4.3 Ablation Study 28 4.3.1 Visual Encoder Size 28 4.3.2 Multi-layer Design 29 4.3.3 Fusion Module Design 30 4.3.4 Text Enhancement Module Design 31 4.3.5 Classification Head Design 32 4.3.6 Loss Function Design 32 4.4 Visualize Result 33 5 Conclusions 37 References 39

    [1] Wing-Sing Cheung, Min-Hsuan You, Si-Yao Syu, Yu-Hsun Chou, and Chi-Yeh Chen. Advancing semantic segmentation of two-dimensional materials using a semantic-adaptive transformer model. Applied Physics Letters, 125(13):133102, 09 2024.
    [2] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.
    [3] Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang. Vision-language transformer and query generation for referring segmentation. In Proceedings of the IEEE International Conference on Computer Vision, 2021.
    [4] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2021.
    [5] Enlai Gao, Shao-Zhen Lin, Zhao Qin, Markus J Buehler, Xi-Qiao Feng, and Zhiping Xu. Mechanical exfoliation of two-dimensional materials. Journal of the Mechanics and Physics of Solids, 115:248–262, 2018.
    [6] Steven M. George. Atomic layer deposition: An overview. Chemical Reviews, pages 111–131, 2010.
    [7] Timothy H. Gfroerer. Photoluminescence in analysis of surfaces and interfaces. Encyclopedia of analytical chemistry, 67:3810, 2000.
    [8] Bingnan Han, Yuxuan Lin, Yafang Yang, Nannan Mao, Wenyue Li, Haozhe Wang, Kenji Yasuda, Xirui Wang, Valla Fatemi, Lin Zhou, Joel I.-Jan Wang, Qiong Ma, Yuan Cao, Daniel Rodan-Legrain, Ya-Qing Bie, Efr´ en Navarro-Moratalla, Dahlia Klein, David MacNeill, Sanfeng Wu, Hikari Kitadai, Xi Ling, Pablo Jarillo-Herrero, Jing Kong, Jihao Yin, and Tom´ as Palacios. Deep-learning-enabled fast optical identification and characterization of 2d materials. Advanced Materials, 32(29):2000953, 2020. 39
    [9] Kaiming He, Georgia Gkioxari, Piotr Doll´ ar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017.
    [10] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
    [11] Ronghang Hu, Marcus Rohrbach, and Trevor Darrell. Segmentation from natural language expressions. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14, pages 108–124. Springer, 2016.
    [12] Shaofei Huang, Tianrui Hui, Si Liu, Guanbin Li, Yunchao Wei, Jizhong Han, Luoqi Liu, and Bo Li. Referring image segmentation via cross-modal progressive compre-hension. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10488–10497, 2020.
    [13] Tianrui Hui, Si Liu, Shaofei Huang, Guanbin Li, Sansi Yu, Faxi Zhang, and Jizhong Han. Linguistic structure guided context modeling for referring image segmentation. In European Conference on Computer Vision, pages 59–75. Springer, 2020.
    [14] Ya Jing, Tao Kong, Wei Wang, Liang Wang, Lei Li, and Tieniu Tan. Locate then segment: A strong pipeline for referring image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9858–9867, 2021.
    [15] Alexander V Kolobov and Junji Tominaga. Two-dimensional transition-metal dichalcogenides, volume 239. Springer, 2016.
    [16] Bart J. Kooi and Beatriz Noheda. Ferroelectric chalcogenides—materials at the edge. Science, 353(6296):221–222, 2016.
    [17] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In F. Pereira, C.J. Burges, L. Bottou, and K.Q. Weinberger, editors, Advances in Neural Information Processing Systems, volume 25. Curran Associates, Inc., 2012. 40
    [18] Ruiyu Li, Kaican Li, Yi-Chun Kuo, Michelle Shu, Xiaojuan Qi, Xiaoyong Shen, and Jiaya Jia. Referring image segmentation via recurrent refinement networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5745–5753, 2018.
    [19] Xiao-Li Li, Wen-Peng Han, Jiang-Bin Wu, Xiao-Fen Qiao, Jun Zhang, and Ping-Heng Tan. Layer-number dependent optical properties of 2d materials and their application for thickness determination. Advanced Functional Materials, 27(19):1604468, 2017.
    [20] Chang Liu, Henghui Ding, and Xudong Jiang. GRES: Generalized referring expression segmentation. In CVPR, 2023.
    [21] Chenxi Liu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, and Alan Yuille. Recurrent multimodal interaction for referring image segmentation. In Proceedings of the IEEE international conference on computer vision, pages 1271–1280, 2017.
    [22] Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021.
    [23] Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11976–11986, 2022.
    [24] Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3431–3440, 2015.
    [25] Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2017.
    [26] Soroush Mahjoubi, Fan Ye, Yi Bao, Weina Meng, and Xian Zhang. Identification and classification of exfoliated graphene flakes from microscopy images using a hierarchical deep convolutional neural network. Engineering Applications of Artificial Intelligence, 119:105743, 2023. 41
    [27] K. S. Novoselov, A. K. Geim, S. V. Morozov, D. Jiang, Y. Zhang, S. V. Dubonos, I. V. Grigorieva, and A. A. Firsov. Electric field effect in atomically thin carbon films. Science, 306(5696):666–669, 2004.
    [28] Ozan Oktay, Jo Schlemper, Lo¨ ıc Le Folgoc, Matthew C. H. Lee, Mattias P. Heinrich, Kazunari Misawa, Kensaku Mori, Steven G. McDonagh, Nils Y. Hammerla, Bernhard Kainz, Ben Glocker, and Daniel Rueckert. Attention u-net: Learning where to look for the pancreas. CoRR, abs/1804.03999, 2018.
    [29] Xuebin Qin, Zichen Zhang, Chenyang Huang, Masood Dehghan, Osmar R. Zaiane, and Martin Jagersand. U2-net: Going deeper with nested u-structure for salient object detection. Pattern Recognition, 106:107404, 2020.
    [30] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pages 234–241. Springer, 2015.
    [31] Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for largescale image recognition. In Yoshua Bengio and Yann LeCun, editors, ICLR, 2014.
    [32] Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmentation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 7262–7272, 2021.
    [33] Si-Yao Syu and Chi-Yeh Chen. Generalization-enhanced transformer integrated with optical microscopy for automated searching and semantic segmentation of two-dimensional materials. Master’s thesis, National Cheng Kung University, Department of Computer Science and Information Engineering, 2024.
    [34] Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning, pages 6105–6114. PMLR, 2019.
    [35] Armin VahidMohammadi, Johanna Rosen, and Yury Gogotsi. The world of two-dimensional carbides and nitrides (mxenes). Science, 372(6547):eabf1581, 2021.42
    [36] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan NGomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
    [37] Qing Hua Wang, Kourosh Kalantar-Zadeh, Andras Kis, Jonathan N Coleman, and Michael S Strano. Electronics and optoelectronics of two-dimensional transition metal dichalcogenides. Nature nanotechnology, 7(11):699–712, 2012.
    [38] Zhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao, Yandong Guo, Mingming Gong, and Tongliang Liu. Cris: Clip-driven referring image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11686–11695, June 2022.
    [39] Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. Transformers: State-of-the-art natural language processing. In Qun Liu and David Schlangen, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online, October 2020. Association for Computational Linguistics.
    [40] Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. Advances in Neural Information Processing Systems, 34:12077–12090, 2021.
    [41] Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, and Philip HS Torr. Lavt: Language-aware vision transformer for referring image segmentation. In CVPR, 2022.
    [42] Min-Hsuan You and Chi-Yeh Chen. Enhancing two-dimensional material semantic segmentation with semantic-adaptive transformer. Master’s thesis, National Cheng Kung University, Department of Computer Science and Information Engineering, 2023.
    [43] Shuang Zhang, Jiong Yang, Renjing Xu, Fan Wang, Weifeng Li, Muhammad Ghufran, Yong Wei Zhang, Zongfu Yu, Gang Zhang, Qinghua Qin, and Yuerui Lu. Extraordinary 43 photoluminescence and strong temperature/angle-dependent raman responses in few layer phosphorene. ACS Nano, 8(9):9590–9596, September 2014.
    [44] Weijie Zhao, Zohreh Ghorannevis, Leiqiang Chu, Minglin Toh, Christian Kloc, PingHeng Tan, and Goki Eda. Evolution of electronic structure in atomically thin sheets of ws2 and wse2. ACS nano, 7(1):791–797, 2013.
    [45] Zongwei Zhou, Md Mahfuzur Rahman Siddiquee, Nima Tajbakhsh, and Jianming Liang. Unet++: A nested u-net architecture for medical image segmentation. In Deep learning in medical image analysis and multimodal learning for clinical decision support: 4th international workshop, DLMIA 2018, and 8th international workshop, ML-CDS 2018, held in conjunction with MICCAI 2018, Granada, Spain, September 20, 2018, proceedings 4, pages 3–11. Springer, 2018.

    下載圖示
    2026-08-31公開
    QR CODE