簡易檢索 / 詳目顯示

研究生: 林峻逸
Lin, Chun-Yi
論文名稱: 基於錯誤模板引導大型語言模型進行構音障礙語音辨識與臨床摘要生成
Error Pattern-Guided Dysarthric Speech Recognition and Clinical Summary Generation Using Large Language Models
指導教授: 吳宗憲
Wu, Chung-Hsien
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 人工智慧科技碩士學位學程
Graduate Program of Artificial Intelligence
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 79
中文關鍵詞: 構音障礙語音辨識錯誤模板音素辨識大型語言模型臨床摘要生成
外文關鍵詞: dysarthric speech recognition, error pattern, phoneme recognition, large language model, Clinical Summary generation
相關次數: 點閱:4下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究提出一套結合錯誤模板(Error Pattern)與大型語言模型之構音障礙語音辨識與臨床摘要生成架構,旨在提升構音障礙語音辨識的準確性,同時提供具可解釋性的臨床診斷資訊。構音障礙語音因受到構音器官控制能力受損影響,常伴隨發音錯誤、韻律異常與高度聲學變異,使得傳統自動語音辨識模型難以有效辨識,且多數方法僅輸出辨識文字或嚴重程度分數,無法描述病患實際的發音偏差與構音特徵。
    為解決上述問題,本研究首先利用 HuBERT-CTC 建立音素辨識模型,預測輸入語音之音素序列,並根據預測音素與參考音素之間的對齊結果建立錯誤模板,描述替換、刪除、插入及正確等構音錯誤資訊。透過錯誤模板預測模型學習構音障礙患者常見之發音偏差模式,並結合音素序列及錯誤模板,引導大型語言模型生成具臨床意義之發音摘要,同時協助恢復病患原始欲表達之詞彙。
    本研究於 UASpeech與Torgo構音障礙語音資料集進行實驗,結果顯示,所提出之方法能有效學習構音錯誤模式,提升構音障礙語音辨識效能,並產生具有可讀性與臨床解釋性的發音摘要,提供病患與家屬更完整的診斷資訊。據我們所知,本研究為首度將錯誤模板預測、臨床摘要生成及大型語言模型整合於構音障礙語音辨識流程,不僅提升語音辨識能力,也建立兼具可解釋性與臨床應用價值之智慧語音分析框架,可望作為未來智慧語言治療與構音障礙評估系統的重要基礎。

    This study proposes an error pattern-guided framework for dysarthric speech recognition and Clinical Summary generation using large language models. The proposed framework aims to improve recognition accuracy while providing interpretable clinical information for speech disorder assessment. Dysarthric speech is characterized by articulatory impairments, abnormal prosody, and high acoustic variability, making it difficult for conventional automatic speech recognition systems to accurately recognize speech. Moreover, most existing approaches only produce recognized transcripts or severity scores, without describing the underlying articulatory errors and speech characteristics.[1]
    To address these challenges, we first employ a HuBERT-CTC model to predict phoneme sequences from dysarthric speech. Error patterns are then derived by aligning the predicted phoneme sequence with the reference sequence, capturing articulatory deviations including substitution, deletion, insertion, and correct phoneme productions. An error pattern detector model is trained to learn common pronunciation deviations in dysarthric speech. The predicted phoneme sequence and error pattern are subsequently integrated into a large language model to generate clinically meaningful pronunciation summaries while simultaneously assisting the recovery of the intended spoken words.
    Experiments conducted on the UASpeech and TORGO dysarthric speech corpora demonstrate that the proposed framework effectively learns articulatory error patterns, improves dysarthric speech recognition performance, and generates readable and clinically interpretable pronunciation summaries. The generated summaries provide patients and their caregivers with more comprehensive diagnostic information. To the best of our knowledge, this is the first work to integrate error pattern detector, Clinical Summary generation, and large language models into a unified dysarthric speech recognition framework. The proposed approach not only enhances recognition performance but also establishes an interpretable and clinically applicable speech analysis framework, offering a promising direction for future intelligent speech therapy and dysarthria assessment systems.

    摘要 I Abstract III 致謝 V Content VII List of Tables X List of Figures XI Chapter 1 Introduction 1 1.1 Background 1 1.2 Problem Statement 3 1.3 Motivation 4 1.4 Literature Review 5 1.4.1 Dysarthric Speech Recognition 5 1.4.2 Pronunciation Error Detection 7 1.4.3 Large Language Models for Clinical Speech Analysis 9 1.5 Proposed Method 10 Chapter 2 Research Methods 13 2.1 Overview of the Proposed Framework 13 2.1.1 Training Framework 13 2.1.2 Inference Framework 14 2.2 Phoneme Recognition Module 16 2.2.1 HuBERT-based Acoustic Encoder 17 2.2.2 CTC-based Phoneme Prediction 18 2.3 Error Pattern Detection Module 19 2.3.1 Error Pattern Generation 19 2.3.2 HuBERT–Phoneme Cross-Attention Detector 22 2.3.3 CTC Training and Error-Pattern Decoding 23 2.4 Clinical Summary Generation 24 2.4.1 Knowledge-Guided Prompt Design 25 2.4.2 Clinical Summary Generation using Gemini 28 2.5 LLM-based Dysarthric Speech Recognition 29 2.5.1 Clinical Summary Generation using Gemma 30 2.5.2 Target Word Prediction using Gemma 32 2.6 Chapter Summary 33 Chapter 3 Experimental Setup and Results 35 3.1 Datasets 35 3.1.1 UASpeech 36 3.1.2 TORGO 38 3.2 Implementation Details 39 3.2.1 HuBERT-based Phoneme Recognition 40 3.2.2 Error Pattern Detector 41 3.2.3 Clinical Summary Generation 42 3.2.4 LLM-based Dysarthric Speech Recognition 43 3.3 Evaluation Metrics 44 3.3.1 Word Error Rate 44 3.3.2 Primary Pattern Accuracy 45 3.3.3 Clinical Summary Human Evaluation 45 3.4 Experimental Results 47 3.4.1 Word Recognition Results 47 3.4.2 Error Pattern Detector Results 51 3.4.3 Clinical Summary Human Evaluation 52 3.5 Ablation Study 54 Chapter 4 Conclusion and Future Work 57 4.1 Conclusion 57 4.2 Future Work 58 Reference 60 Appendix A. Prompt Templates 63 A.1 Clinical Summary Generation Prompt (Gemini) 63 A.1.1 Instruction Prompt 63 A.2 Target Word Prediction Prompt (Gemma) 64 A.2.1 System Prompt 64 A.2.2 User Prompt 65 A.3 Output Constraints 65

    [1] K. Singhal et al., "Large language models encode clinical knowledge," Nature, vol. 620, no. 7972, pp. 172–180, 2023/08/01 2023, doi: 10.1038/s41586-023-06291-2.
    [2] R. D. Kent, G. Weismer, J. F. Kent, H. K. Vorperian, and J. R. Duffy, "Acoustic studies of dysarthric speech: methods, progress, and potential," (in eng), J Commun Disord, vol. 32, no. 3, pp. 141–80, 183–6; quiz 181–3, 187–9, May–Jun 1999, doi: 10.1016/s0021-9924(99)00004-0.
    [3] M. Tu, A. Wisler, V. Berisha, and J. M. Liss, "The relationship between perceptual disturbances in dysarthric speech and automatic speech recognition performance," (in eng), J Acoust Soc Am, vol. 140, no. 5, p. El416, Nov 2016, doi: 10.1121/1.4967208.
    [4] A. Baevski, H. Zhou, A. Mohamed, and M. Auli, "wav2vec 2.0: a framework for self-supervised learning of speech representations," presented at the Proceedings of the 34th International Conference on Neural Information Processing Systems, Vancouver, BC, Canada, 2020.
    [5] W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, "Hubert: Self-supervised speech representation learning by masked prediction of hidden units," IEEE/ACM transactions on audio, speech, and language processing, vol. 29, pp. 3451–3460, 2021.
    [6] C. Bhat and H. Strik, "Speech Technology for Automatic Recognition and Assessment of Dysarthric Speech: An Overview," (in eng), J Speech Lang Hear Res, vol. 68, no. 2, pp. 547–577, Feb 4 2025, doi: 10.1044/2024_jslhr-23-00740.
    [7] R. D. Kent and Y. J. Kim, "Toward an acoustic typology of motor speech disorders," Clinical linguistics & phonetics, vol. 17, no. 6, pp. 427–445, 2003.
    [8] V. Wolfrum, K. Lehner, S. Heim, and W. Ziegler, "Clinical Assessment of Communication-Related Speech Parameters in Dysarthria: The Impact of Perceptual Adaptation," (in eng), J Speech Lang Hear Res, vol. 66, no. 8, pp. 2622–2642, Aug 3 2023, doi: 10.1044/2023_jslhr-23-00105.
    [9] Z. Qian, K. Xiao, and C. Yu, "A survey of technologies for automatic Dysarthric speech recognition," EURASIP Journal on Audio, Speech, and Music Processing, vol. 2023, 11/11 2023, doi: 10.1186/s13636-023-00318-2.
    [10] N. Hiebel, O. Ferret, K. Fort, and A. Névéol, "Clinical Text Generation: Are We There Yet?," (in eng), Annu Rev Biomed Data Sci, vol. 8, no. 1, pp. 173–198, Aug 2025, doi: 10.1146/annurev-biodatasci-103123-095202.
    [11] J. R. Deller, Jr., D. Hsu, and L. J. Ferrier, "On the use of hidden Markov modelling for recognition of dysarthric speech," (in eng), Comput Methods Programs Biomed, vol. 35, no. 2, pp. 125–39, Jun 1991, doi: 10.1016/0169-2607(91)90071-z.
    [12] S. Chen et al., "WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing," IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022, doi: 10.1109/JSTSP.2022.3188113.
    [13] S. Wang, S. Zhao, J. Zhou, A. Kong, and Y. Qin, "Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation," in Interspeech 2024, 2024, pp. 1305–1309, doi: 10.21437/Interspeech.2024-1360. [Online]. Available: https://doi.org/10.21437/Interspeech.2024-1360
    [14] W. Lee, S. Im, H. Do, Y. Kim, J. Ok, and G. Lee, "DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition," Albuquerque, New Mexico, April 2025: Association for Computational Linguistics, in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 4701–4712, doi: 10.18653/v1/2025.naacl-long.240. [Online]. Available: https://aclanthology.org/2025.naacl-long.240/
    [15] I. T. Hsieh and C. H. Wu, "Hierarchical Curriculum Learning for Dysarthric Speech Recognition via Multi-Level Knowledge Distillation," IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 3553–3567, 2025, doi: 10.1109/TASLPRO.2025.3597438.
    [16] K. Fu, J. Lin, D. Ke, Y. Xie, J. Zhang, and B. Lin, A Full Text-Dependent End to End Mispronunciation Detection and Diagnosis with Easy Data Augmentation Techniques. 2021.
    [17] C. Zhu et al., "Pronunciation error detection model based on feature fusion," Speech Communication, vol. 156, p. 103009, 2024/01/01/ 2024, doi: https://doi.org/10.1016/j.specom.2023.103009.
    [18] X. Zhou et al., Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling. 2025, pp. 4738–4742.
    [19] J. Wu, X. Wu, and J. Yang, "Guiding clinical reasoning with large language models via knowledge seeds," presented at the Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, Jeju, Korea, 2024. [Online]. Available: https://doi.org/10.24963/ijcai.2024/829.
    [20] N. Kim, M. Homer, and H. Jang, "Clinical Application of Large Language Models for Intervention Plan Development in Speech-Language Pathology," (in eng), Am J Speech Lang Pathol, vol. 34, no. 4, pp. 2098–2114, Jul 10 2025, doi: 10.1044/2025_ajslp-24-00464.
    [21] T. Arias-Vergara et al., Acoustic-Driven Generation of Pathological Speech Reports Using Large Language Models. 2025.
    [22] S. Park, C. Gupta, M. Kwan, X. Fung, A. Yip, and S. Nanayakkara, Towards Temporally Explainable Dysarthric Speech Clarity Assessment. 2025.
    [23] A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural 'networks. 2006, pp. 369–376.
    [24] V. I. Levenshtein, "Binary codes capable of correcting deletions, insertions, and reversals," Soviet physics. Doklady, vol. 10, pp. 707–710, 1965.
    [25] S. Hochreiter and J. Schmidhuber, "Long Short-Term Memory," Neural Computation, vol. 9, pp. 1735–1780, 11/15 1997, doi: 10.1162/neco.1997.9.8.1735.
    [26] A. Vaswani et al., "Attention is all you need," presented at the Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, California, USA, 2017.
    [27] H. Kim et al., "Dysarthric speech database for universal access research," in Interspeech, 2008, vol. 2008, pp. 1741–1744.
    [28] F. Rudzicz, A. K. Namasivayam, and T. Wolff, "The TORGO database of acoustic and articulatory speech from speakers with dysarthria," Language resources and evaluation, vol. 46, pp. 523–541, 2012.

    QR CODE