簡易檢索 / 詳目顯示

研究生: 劉哲嘉
LIU, CHE-CHIA
論文名稱: 透過自動提示引導多任務模型應用於定位胸腔X 光片上的氣管導管與氣管隆突
Automatic Prompt-based Multi-task Model for Endotracheal Tube and Carina Landmark Detection in Portable Supine Chest Radiographs
指導教授: 陳奇業
Chen, Chi-Yeh
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 醫學資訊研究所
Institute of Medical Informatics
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 59
中文關鍵詞: 深度學習多任務學習氣管插管實例分割關鍵點偵測
外文關鍵詞: Deep learning, Multi-task learning, Endotracheal intubation, Instance segmentation, Landmark detection
相關次數: 點閱:10下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 氣管插管作為緊急醫療程序,在加護病房中當病人無法自主呼吸時被廣泛使用,在執行插管後必須確定氣管導管的位置是否適當,以降低併發症的風險。然而,在加護病房中,常使用可攜式的X光機搭配仰臥位的前後X光視角 (supine AP view) 來評估氣管導管位置,這種影像獲取方式常遇到遮擋以及外部設備的干擾,且整體影像品質可能較差,在自動化定位導管末端以及氣管隆突的任務上面臨挑戰。本論文提出了一個端到端的多任務學習模型,該模型能夠同時進行實例分割以及氣管導管末端和氣管隆突的關鍵點偵測。該模型採用提示引導架構,然而一般提示引導架構中,需要人工提供提示,因此本論文透過自動提示生成模組 (Automatic Prompt Generation) 完全免除人工提示的需求,使用預先生成的粗略分割遮罩來產生幾何提示例如:點、物件框,並將這些提示結合到模型中以引導精確的定位。此外,我們還開發了邊緣引導的遮罩編碼器 (Edge-Guided Mask Encoder)、跨任務注意力模組 (Cross-Task Token Attention)、任務決定提示注意力模組 (Task-Conditioned Prompt Attention) 以及引入多任務學習損失函數平衡分割與關鍵點偵測的訓練。
    在國立成功大學醫學院附屬醫院所提供的資料集上進行實驗,於預測導管錯位的評估中,本論文提出的模型在內部驗證中達到 $90.81%$ 的準確率,預測的導管末端到氣管隆突的距離平均誤差為 $4.96 pm 6.33$ mm ,在跨機構的外部測試集上具備 $90.00%$ 的準確率以及 $4.78 pm 6.00$ mm 的距離平均誤差。在關鍵點偵測評估中,模型對於導管末端的位置平均誤差為 $3.68 pm 4.93$ mm,對於氣管隆突的位置平均誤差為 $3.79 pm 3.56$ mm。在分割評估中,端到端的模型達到 $77.19%$ 的平均 Dice 係數和 $64.14%$ 的平均 IoU。這些結果表明,該模型在導管位置評估方面超越了過去的方法,且同時具備關鍵點偵測與分割的能力。

    Accurate monitoring of endotracheal tube (ETT) position is essential in intensive care units, yet automated localization on portable chest radiographs remains challenging because of overlapping anatomical structures, low contrast, and interference from external devices. This thesis presents an end-to-end, prompt-based multi-task framework that jointly performs instance segmentation and landmark detection for the ETT tip and carina. To eliminate manual prompt interaction, we introduce an Automatic Prompt Generation module that converts coarse segmentation masks into geometric prompts. To improve boundary-aware representation learning and reduce task interference, we further incorporate an Edge-Guided Mask Encoder, Cross-Task Token Attention, and Task-Conditioned Prompt Attention. In addition, we adopt an uncertainty-driven multi-task loss to balance segmentation and landmark detection during stable joint training.
    Experiments were conducted on a dataset from National Cheng Kung University Hospital and Tainan Hospital. For ETT depth estimation, the proposed framework achieves 90.81% accuracy and a mean absolute error (MAE) of Tip-Carina distance of 4.96 ± 6.33 mm in 5-fold cross-validation. On the external test set, it maintains 90.00% accuracy and an MAE of 4.78 ± 6.00 mm, while producing zero failure cases in both internal and external evaluations. For landmark detection, the model achieves external mean radial errors of 3.68 ± 4.93 mm for the ETT tip and 3.79 ± 3.56 mm for the carina. For segmentation, the fully automated auto prompt generation setting achieves 77.19% mDice and 64.14% mIoU. These results demonstrate that the proposed framework provides robust, end-to-end, and clinically viable support for endotracheal tube and carina landmark detection and segmentation.

    中文摘要 i Abstract iii 誌謝 v Contents vi List of Tables viii List of Figures ix 1 Introduction 1 1.1 Overview 1 1.2 Motivation 2 1.3 Contribution 3 2 Related Work 5 2.1 Foundation Models for Medical Image Segmentation 5 2.1.1 Vision Transformer for Medical Segmentation 5 2.1.2 Segment Anything Model and Medical Adaptations 6 2.1.3 Automatic Prompt Generation 7 2.2 Landmark Detection in Medical Imaging 7 2.2.1 Detection and Heatmap Regression Approaches 8 2.2.2 Transformer-based Approaches 8 2.3 Multi-task Learning for Segmentation and Landmark Detection 9 3 Methods 10 3.1 Overview 10 3.2 Image Encoder 11 3.3 Automatic Prompt Generation 12 3.4 Prompt Encoder 13 3.5 Multi-task Decoder 16 3.5.1 Two-Way Transformer with Cross-Task Token Attention 16 3.5.2 Task-Conditioned Prompt Attention 19 3.5.3 Segmentation Branch 21 3.5.4 Landmark Detection Branch 21 3.6 Loss Function 23 4 Experiments 25 4.1 Dataset 25 4.2 Evaluation Metrics 26 4.3 Implementation Details 28 4.4 Experimental Results 28 4.5 Comparison with State-of-the-Art Methods 30 4.6 Comparison with Segmentation Methods 32 4.7 Ablation Study 34 4.8 Visualization Results 36 5 Conclusion 39 References 41

    [1] Liang-Kai Mao, Min-Hsin Huang, Chao-Han Lai, Yung-Nien Sun, and Chi-Yeh Chen, “Detecting endotracheal tube and carina on portable supine chest radiographs using one-stage detector with a coarse-to-fine attention,” Diagnostics, vol. 12, no. 8, 2022, ISSN: 2075-4418.
    [2] Min-Hsin Huang, Chi-Yeh Chen, M. Horng, Chung-I Li, I-Lin Hsu, Che-Min Su, Yung-Nien Sun, and Chao-Han Lai, “Validation of a deep learning–based automatic detection algorithm for measurement of endotracheal tube–to–carina distance on chest radiographs,” Anesthesiology, vol. 137, pp. 704–715, 2022.
    [3] Manu Varshney, Kavita Sharma, Rakesh Kumar, and Preeti G Varshney, “Appropriate depth of placement of oral endotracheal tube and its possible determinants in indian adult patients,” Indian Journal of Anaesthesia, vol. 55, no. 5, pp. 488–493, 2011.
    [4] P. Lakhani, A. Flanders, and R. Gorniak, “Endotracheal tube position assessment on chest radiographs using deep learning.,” Radiology. Artificial intelligence, vol. 3 1, e200026, 2020.
    [5] Maayan Frid-Adar, Rula Amer, and Hayit Greenspan, “Endotracheal tube detection and segmentation in chest radiographs using synthetic data,” International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer, 2019, pp. 784–792.
    [6] Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick, “Mask r-cnn,” Proceedings of the IEEE international conference on computer vision, 2017, pp. 2961–2969.
    [7] Zhi Tian, Chunhua Shen, Hao Chen, and Tong He, “Fcos: Fully convolutional onestage object detection,” Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9627–9636.
    [8] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” International Conference on Learning Representations, 2021.
    [9] Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao, “ViTPose: Simple vision transformer baselines for human pose estimation,” Advances in Neural Information Processing Systems, 2022.
    [10] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick, “Segment anything,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2023, pp. 4015–4026.
    [11] Chaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei, Haoqi Fan, Po-Yao Huang, Vaibhav Aggarwal, Arkabandhu Chowdhury, Omid Poursaeed, Judy Hoffman, Jitendra Malik, Yanghao Li, and Christoph Feichtenhofer, “Hiera: A hierarchical vision transformer without the bells-and-whistles,” Proceedings of the 40th International Conference on Machine Learning, ser. ICML’23, Honolulu, Hawaii, USA: JMLR.org, 2023.
    [12] Jun Ma, Yuting He, Feifei Li, Lin Han, Chenyu You, and Bo Wang, “Segment anything in medical images,” Nature Communications, vol. 15, no. 1, Jan. 2024, ISSN: 20411723.
    [13] Yichi Zhang, Zhenrong Shen, and Rushi Jiao, “Segment anything model for medical image segmentation: Current applications and future directions,” Computers in Biology and Medicine, vol. 171, p. 108 238, 2024.
    [14] Junde Wu, Ziyue Wang, Mingxuan Hong, Wei Ji, Huazhu Fu, Yanwu Xu, Min Xu, and Yueming Jin, “Medical sam adapter: Adapting segment anything model for medical image segmentation,” Medical image analysis, vol. 102, p. 103 547, 2025.
    [15] Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollar, and Christoph Feichtenhofer, “SAM 2: Segment anything in images and videos,” The Thirteenth International Conference on Learning Representations, 2025.
    [16] Xinyu Xiong, Zihuang Wu, Shuangyi Tan, Wenxue Li, Feilong Tang, Ying Chen, Siying Li, Jie Ma, and Guanbin Li, “Sam2-unet: Segment anything 2 makes strong encoder for natural and medical image segmentation,” Visual Intelligence, vol. 4, no. 1, p. 2, 2026.
    [17] Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 9992–10 002.
    [18] Jeya Maria Jose Valanarasu, Poojan Oza, Ilker Hacihaliloglu, and Vishal M Patel, “Medical transformer: Gated axial-attention for medical image segmentation,” International conference on medical image computing and computer-assisted intervention, Springer, 2021, pp. 36–46.
    [19] Jieneng Chen, Jieru Mei, Xianhang Li, Yongyi Lu, Qihang Yu, Qingyue Wei, Xiangde Luo, Yutong Xie, Ehsan Adeli, Yan Wang, Matthew P. Lungren, Shaoting Zhang, Lei Xing, Le Lu, Alan Yuille, and Yuyin Zhou, “Transunet: Rethinking the u-net architecture design for medical image segmentation through the lens of transformers,” Medical Image Analysis, vol. 97, p. 103 280, 2024, ISSN: 1361-8415.
    [20] Ali Hatamizadeh, Dong Yang, H. Roth, and Daguang Xu, “UNETR: Transformers for 3d medical image segmentation,” 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), vol. null, pp. 1748–1758, 2021.
    [21] Ali Hatamizadeh, Vishwesh Nath, Yucheng Tang, Dong Yang, Holger R. Roth, and Daguang Xu, “Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images,” Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries, Alessandro Crimi and Spyridon Bakas, Eds., Cham: Springer International Publishing, 2022, pp. 272–284, ISBN: 978-3-031-08999-2.
    [22] Shizhan Gong, Yuan Zhong, Wenao Ma, Jinpeng Li, Zhao Wang, Jingyang Zhang, Pheng-Ann Heng, and Qi Dou, “3dsam-adapter: Holistic adaptation of sam from 2d to 3d for promptable tumor segmentation,” Medical Image Analysis, vol. 98, p. 103 324, 2024.
    [23] Chengyin Li, Prashant Khanduri, Yao Qiang, Rafi Ibn Sultan, Indrin J. Chetty, and D. Zhu, “Autoprosam: Automated prompting sam for 3d multi-organ segmentation,” 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 3570–3580, 2023.
    [24] Saiyang Na, Yuzhi Guo, Feng Jiang, Hehuan Ma, Jean Gao, and Junzhou Huang, “Segment any cell: A sam-based auto-prompting fine-tuning framework for nuclei segmentation,” IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 12, pp. 19 986–19 995, 2025.
    [25] Danping Yin, Qingqing Zheng, Long Chen, Ying Hu, and Qiong Wang, “Apg-sam: Automatic prompt generation for sam-based breast lesion segmentation with boundaryaware optimization,” Expert Systems with Applications, vol. 276, p. 127 048, 2025.
    [26] Ching-Wei Wang, Cheng-Ta Huang, Meng-Che Hsieh, Chung-Hsing Li, Sheng-Wei Chang, Wei-Cheng Li, Rémy Vandaele, Raphaël Marée, Sébastien Jodogne, Pierre Geurts, Cheng Chen, Guoyan Zheng, Chengwen Chu, Hengameh Mirzaalian, Ghassan Hamarneh, Tomaž Vrtovec, and Bulat Ibragimov, “Evaluation and comparison of anatomical landmark detection methods for cephalometric x-ray images: A grand challenge,” IEEE Transactions on Medical Imaging, vol. 34, no. 9, pp. 1890–1900, 2015.
    [27] Runnan Chen, Yuexin Ma, Nenglun Chen, Lingjie Liu, Zhiming Cui, Yanhong Lin, and Wenping Wang, “Structure-aware long short-term memory network for 3d cephalometric landmark detection,” IEEE Transactions on Medical Imaging, vol. 41, no. 7, pp. 1791–1801, 2022.
    [28] Jun Zhang, Mingxia Liu, and Dinggang Shen, “Detecting anatomical landmarks from limited medical imaging data using two-stage task-oriented deep neural networks,” IEEE Transactions on Image Processing, vol. 26, no. 10, pp. 4753–4764, 2017.
    [29] Yankun Lang, Chunfeng Lian, Deqiang Xiao, Hannah Deng, Kim-Han Thung, Peng Yuan, Jaime Gateno, Tianshu Kuang, David M. Alfi, Li Wang, Dinggang Shen, James J. Xia, and Pew-Thian Yap, “Localization of craniomaxillofacial landmarks on cbct images using 3d mask r-cnn and local dependency learning,” IEEE Transactions on Medical Imaging, vol. 41, no. 10, pp. 2856–2866, 2022.
    [30] Kaiwen Wan, Lei Li, Dengqiang Jia, Shangqi Gao, Wei Qian, Yingzhi Wu, Huandong Lin, Xiongzheng Mu, Xin Gao, Sijia Wang, et al., “Multi-target landmark detection with incomplete images via reinforcement learning and shape prior embedding,” Medical Image Analysis, vol. 89, p. 102 875, 2023.
    [31] Christian Payer, Darko Štern, Horst Bischof, and Martin Urschler, “Regressing heatmaps for multiple landmark localization using cnns,” International conference on medical image computing and computer-assisted intervention, Springer, 2016, pp. 230–238.
    [32] Xiaoyang Chen, Chunfeng Lian, Hannah H. Deng, Tianshu Kuang, Hung-Ying Lin, Deqiang Xiao, Jaime Gateno, Dinggang Shen, James J. Xia, and Pew-Thian Yap, “Fast and accurate craniomaxillofacial landmark detection via 3d faster r-cnn,” IEEE Transactions on Medical Imaging, vol. 40, no. 12, pp. 3867–3878, 2021.
    [33] Zixun Huang, Rui Zhao, Frank H. F. Leung, Sunetra Banerjee, Kin-Man Lam, YongPing Zheng, and Sai Ho Ling, “Landmark localization from medical images with generative distribution prior,” IEEE Transactions on Medical Imaging, vol. 43, no. 7, pp. 2679–2692, 2024.
    [34] Alexandra Ertl, Stefan Denner, Robin Peretzke, Shuhan Xiao, David Zimmerer, Maximilian Fischer, Markus Ralf Bujotzek, Xin Yang, Peter Neher, Fabian Isensee, and Klaus Maier-Hein, “Nnlandmark: A self-configuring method for 3d medical landmark detection,” Medical Imaging with Deep Learning, 2026.
    [35] Fabian Isensee, Paul F Jaeger, Simon AA Kohl, Jens Petersen, and Klaus H MaierHein, “Nnu-net: A self-configuring method for deep learning-based biomedical image segmentation,” Nature methods, vol. 18, no. 2, pp. 203–211, 2021.
    [36] Yanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang, Wankou Yang, Shu-Tao Xia, and Erjin Zhou, “Tokenpose: Learning keypoint tokens for human pose estimation,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021, pp. 11 313–11 322.
    [37] Yan Zhao, Xiuying Wang, Tongtong Che, Guoqing Bao, and Shuyu Li, “Multi-task deep learning for medical image computing and analysis: A review,” Computers in Biology and Medicine, vol. 153, p. 106 496, 2023.
    [38] Zimeng Tan, Jianjiang Feng, and Jie Zhou, “Multi-task learning network for landmark detection in anatomical tree structures,” 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), IEEE, 2021, pp. 1975–1979.
    [39] Dong Miao, Ying Zhao, Xue Ren, Meng Dou, Yu Yao, Yiran Xu, Yingchao Cui, and Ailian Liu, “A multi-task based deep learning framework with landmark detection for mri couinaud segmentation,” IEEE Journal of Translational Engineering in Health and Medicine, vol. 12, pp. 697–710, 2024.
    [40] Xuehao Wang, Zhan Zhuang, Feiyang Ye, and Yu Zhang, “Mtsam: Multi-task finetuning for segment anything model,” The Thirteenth International Conference on Learning Representations, 2025.
    [41] Songtao Liu, Di Huang, et al., “Receptive field block net for accurate and fast object detection,” Proceedings of the European conference on computer vision (ECCV), 2018, pp. 385–400.
    [42] Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng, “Fourier features let networks learn high frequency functions in low dimensional domains,” Proceedings of the 34th International Conference on Neural Information Processing Systems, ser. NIPS ’20, Vancouver, BC, Canada: Curran Associates Inc., 2020, ISBN: 9781713829546.
    [43] Qingming Huang Jun Wei Shuhui Wang, “F3net: Fusion, feedback and focus for salient object detection,” AAAI Conference on Artificial Intelligence (AAAI), 2020.
    [44] Ivan Lopes, Tuan-Hung Vu, and Raoul De Charette, “Cross-task attention mechanism for dense multi-task learning,” Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 2329–2338.
    [45] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017.
    [46] Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu, “Conditional prompt learning for vision-language models,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 16 816–16 825.
    [47] Olaf Ronneberger, Philipp Fischer, and Thomas Brox, “U-net: Convolutional networks for biomedical image segmentation,” International Conference on Medical image computing and computer-assisted intervention, Springer, 2015, pp. 234–241.
    [48] Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville, “Film: Visual reasoning with a general conditioning layer,” Proceedings of the AAAI conference on artificial intelligence, vol. 32, 2018.
    [49] Xinyao Wang, Liefeng Bo, and Li Fuxin, “Adaptive wing loss for robust face alignment via heatmap regression,” The IEEE International Conference on Computer Vision (ICCV), Oct. 2019.
    [50] Alex Kendall, Yarin Gal, and Roberto Cipolla, “Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,” Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7482–7491.
    [51] Ilya Loshchilov and Frank Hutter, “Decoupled weight decay regularization,” International Conference on Learning Representations, 2019.

    QR CODE