| 研究生: |
林峻逸 Lin, Chun-Yi |
|---|---|
| 論文名稱: |
基於錯誤模板引導大型語言模型進行構音障礙語音辨識與臨床摘要生成 Error Pattern-Guided Dysarthric Speech Recognition and Clinical Summary Generation Using Large Language Models |
| 指導教授: |
吳宗憲
Wu, Chung-Hsien |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 人工智慧科技碩士學位學程 Graduate Program of Artificial Intelligence |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 79 |
| 中文關鍵詞: | 構音障礙語音辨識 、錯誤模板 、音素辨識 、大型語言模型 、臨床摘要生成 |
| 外文關鍵詞: | dysarthric speech recognition, error pattern, phoneme recognition, large language model, Clinical Summary generation |
| 相關次數: | 點閱:4 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究提出一套結合錯誤模板(Error Pattern)與大型語言模型之構音障礙語音辨識與臨床摘要生成架構,旨在提升構音障礙語音辨識的準確性,同時提供具可解釋性的臨床診斷資訊。構音障礙語音因受到構音器官控制能力受損影響,常伴隨發音錯誤、韻律異常與高度聲學變異,使得傳統自動語音辨識模型難以有效辨識,且多數方法僅輸出辨識文字或嚴重程度分數,無法描述病患實際的發音偏差與構音特徵。
為解決上述問題,本研究首先利用 HuBERT-CTC 建立音素辨識模型,預測輸入語音之音素序列,並根據預測音素與參考音素之間的對齊結果建立錯誤模板,描述替換、刪除、插入及正確等構音錯誤資訊。透過錯誤模板預測模型學習構音障礙患者常見之發音偏差模式,並結合音素序列及錯誤模板,引導大型語言模型生成具臨床意義之發音摘要,同時協助恢復病患原始欲表達之詞彙。
本研究於 UASpeech與Torgo構音障礙語音資料集進行實驗,結果顯示,所提出之方法能有效學習構音錯誤模式,提升構音障礙語音辨識效能,並產生具有可讀性與臨床解釋性的發音摘要,提供病患與家屬更完整的診斷資訊。據我們所知,本研究為首度將錯誤模板預測、臨床摘要生成及大型語言模型整合於構音障礙語音辨識流程,不僅提升語音辨識能力,也建立兼具可解釋性與臨床應用價值之智慧語音分析框架,可望作為未來智慧語言治療與構音障礙評估系統的重要基礎。
This study proposes an error pattern-guided framework for dysarthric speech recognition and Clinical Summary generation using large language models. The proposed framework aims to improve recognition accuracy while providing interpretable clinical information for speech disorder assessment. Dysarthric speech is characterized by articulatory impairments, abnormal prosody, and high acoustic variability, making it difficult for conventional automatic speech recognition systems to accurately recognize speech. Moreover, most existing approaches only produce recognized transcripts or severity scores, without describing the underlying articulatory errors and speech characteristics.[1]
To address these challenges, we first employ a HuBERT-CTC model to predict phoneme sequences from dysarthric speech. Error patterns are then derived by aligning the predicted phoneme sequence with the reference sequence, capturing articulatory deviations including substitution, deletion, insertion, and correct phoneme productions. An error pattern detector model is trained to learn common pronunciation deviations in dysarthric speech. The predicted phoneme sequence and error pattern are subsequently integrated into a large language model to generate clinically meaningful pronunciation summaries while simultaneously assisting the recovery of the intended spoken words.
Experiments conducted on the UASpeech and TORGO dysarthric speech corpora demonstrate that the proposed framework effectively learns articulatory error patterns, improves dysarthric speech recognition performance, and generates readable and clinically interpretable pronunciation summaries. The generated summaries provide patients and their caregivers with more comprehensive diagnostic information. To the best of our knowledge, this is the first work to integrate error pattern detector, Clinical Summary generation, and large language models into a unified dysarthric speech recognition framework. The proposed approach not only enhances recognition performance but also establishes an interpretable and clinically applicable speech analysis framework, offering a promising direction for future intelligent speech therapy and dysarthria assessment systems.
[1] K. Singhal et al., "Large language models encode clinical knowledge," Nature, vol. 620, no. 7972, pp. 172–180, 2023/08/01 2023, doi: 10.1038/s41586-023-06291-2.
[2] R. D. Kent, G. Weismer, J. F. Kent, H. K. Vorperian, and J. R. Duffy, "Acoustic studies of dysarthric speech: methods, progress, and potential," (in eng), J Commun Disord, vol. 32, no. 3, pp. 141–80, 183–6; quiz 181–3, 187–9, May–Jun 1999, doi: 10.1016/s0021-9924(99)00004-0.
[3] M. Tu, A. Wisler, V. Berisha, and J. M. Liss, "The relationship between perceptual disturbances in dysarthric speech and automatic speech recognition performance," (in eng), J Acoust Soc Am, vol. 140, no. 5, p. El416, Nov 2016, doi: 10.1121/1.4967208.
[4] A. Baevski, H. Zhou, A. Mohamed, and M. Auli, "wav2vec 2.0: a framework for self-supervised learning of speech representations," presented at the Proceedings of the 34th International Conference on Neural Information Processing Systems, Vancouver, BC, Canada, 2020.
[5] W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, "Hubert: Self-supervised speech representation learning by masked prediction of hidden units," IEEE/ACM transactions on audio, speech, and language processing, vol. 29, pp. 3451–3460, 2021.
[6] C. Bhat and H. Strik, "Speech Technology for Automatic Recognition and Assessment of Dysarthric Speech: An Overview," (in eng), J Speech Lang Hear Res, vol. 68, no. 2, pp. 547–577, Feb 4 2025, doi: 10.1044/2024_jslhr-23-00740.
[7] R. D. Kent and Y. J. Kim, "Toward an acoustic typology of motor speech disorders," Clinical linguistics & phonetics, vol. 17, no. 6, pp. 427–445, 2003.
[8] V. Wolfrum, K. Lehner, S. Heim, and W. Ziegler, "Clinical Assessment of Communication-Related Speech Parameters in Dysarthria: The Impact of Perceptual Adaptation," (in eng), J Speech Lang Hear Res, vol. 66, no. 8, pp. 2622–2642, Aug 3 2023, doi: 10.1044/2023_jslhr-23-00105.
[9] Z. Qian, K. Xiao, and C. Yu, "A survey of technologies for automatic Dysarthric speech recognition," EURASIP Journal on Audio, Speech, and Music Processing, vol. 2023, 11/11 2023, doi: 10.1186/s13636-023-00318-2.
[10] N. Hiebel, O. Ferret, K. Fort, and A. Névéol, "Clinical Text Generation: Are We There Yet?," (in eng), Annu Rev Biomed Data Sci, vol. 8, no. 1, pp. 173–198, Aug 2025, doi: 10.1146/annurev-biodatasci-103123-095202.
[11] J. R. Deller, Jr., D. Hsu, and L. J. Ferrier, "On the use of hidden Markov modelling for recognition of dysarthric speech," (in eng), Comput Methods Programs Biomed, vol. 35, no. 2, pp. 125–39, Jun 1991, doi: 10.1016/0169-2607(91)90071-z.
[12] S. Chen et al., "WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing," IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022, doi: 10.1109/JSTSP.2022.3188113.
[13] S. Wang, S. Zhao, J. Zhou, A. Kong, and Y. Qin, "Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation," in Interspeech 2024, 2024, pp. 1305–1309, doi: 10.21437/Interspeech.2024-1360. [Online]. Available: https://doi.org/10.21437/Interspeech.2024-1360
[14] W. Lee, S. Im, H. Do, Y. Kim, J. Ok, and G. Lee, "DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition," Albuquerque, New Mexico, April 2025: Association for Computational Linguistics, in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 4701–4712, doi: 10.18653/v1/2025.naacl-long.240. [Online]. Available: https://aclanthology.org/2025.naacl-long.240/
[15] I. T. Hsieh and C. H. Wu, "Hierarchical Curriculum Learning for Dysarthric Speech Recognition via Multi-Level Knowledge Distillation," IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 3553–3567, 2025, doi: 10.1109/TASLPRO.2025.3597438.
[16] K. Fu, J. Lin, D. Ke, Y. Xie, J. Zhang, and B. Lin, A Full Text-Dependent End to End Mispronunciation Detection and Diagnosis with Easy Data Augmentation Techniques. 2021.
[17] C. Zhu et al., "Pronunciation error detection model based on feature fusion," Speech Communication, vol. 156, p. 103009, 2024/01/01/ 2024, doi: https://doi.org/10.1016/j.specom.2023.103009.
[18] X. Zhou et al., Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling. 2025, pp. 4738–4742.
[19] J. Wu, X. Wu, and J. Yang, "Guiding clinical reasoning with large language models via knowledge seeds," presented at the Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, Jeju, Korea, 2024. [Online]. Available: https://doi.org/10.24963/ijcai.2024/829.
[20] N. Kim, M. Homer, and H. Jang, "Clinical Application of Large Language Models for Intervention Plan Development in Speech-Language Pathology," (in eng), Am J Speech Lang Pathol, vol. 34, no. 4, pp. 2098–2114, Jul 10 2025, doi: 10.1044/2025_ajslp-24-00464.
[21] T. Arias-Vergara et al., Acoustic-Driven Generation of Pathological Speech Reports Using Large Language Models. 2025.
[22] S. Park, C. Gupta, M. Kwan, X. Fung, A. Yip, and S. Nanayakkara, Towards Temporally Explainable Dysarthric Speech Clarity Assessment. 2025.
[23] A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural 'networks. 2006, pp. 369–376.
[24] V. I. Levenshtein, "Binary codes capable of correcting deletions, insertions, and reversals," Soviet physics. Doklady, vol. 10, pp. 707–710, 1965.
[25] S. Hochreiter and J. Schmidhuber, "Long Short-Term Memory," Neural Computation, vol. 9, pp. 1735–1780, 11/15 1997, doi: 10.1162/neco.1997.9.8.1735.
[26] A. Vaswani et al., "Attention is all you need," presented at the Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, California, USA, 2017.
[27] H. Kim et al., "Dysarthric speech database for universal access research," in Interspeech, 2008, vol. 2008, pp. 1741–1744.
[28] F. Rudzicz, A. K. Namasivayam, and T. Wolff, "The TORGO database of acoustic and articulatory speech from speakers with dysarthria," Language resources and evaluation, vol. 46, pp. 523–541, 2012.