簡易檢索 / 詳目顯示

研究生: 潘公福
Phan, Cong-Phuoc
論文名稱: 結合學習方法之統計資料精煉技術於生成式人工智慧模型調適之研究
Statistical Data Refinement with Learning Methods for the Adaptation of Generative Artificial Intelligence
指導教授: 蔣榮先
Chiang, Jung-Hsien
學位類別: 博士
Doctor
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 91
中文關鍵詞: 生成式人工智慧資料填補統計精煉
外文關鍵詞: Generative AI learning methods, Data Imputation, Statistical refinement
相關次數: 點閱:26下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本論文旨在提升人工智慧模型在不同使用者資料領域中的知識能力。此人工智慧模型可以是基於其他資料集進行預訓練的深度學習模型,例如大型語言模型(LLM)、生成式模型,或用於嵌入表示的預訓練模型。透過本框架進行改進後,升級後的模型能夠在特定下游任務中展現顯著優異的效能,包括命名實體辨識(NER)、關係抽取(RE)、情緒偵測/預測,以及針對特定領域需求設計的 Text-to-SQL 轉換任務。
    隨著人工智慧從通用能力逐漸轉向專業化的實際應用場景,僅單純擴大模型規模並不能保證其在專業領域應用中的有效性,這形成了一項關鍵挑戰。為了最大化人工智慧系統對終端使用者的效益,系統必須有效連結原始領域資料輸入與複雜的生成式能力。達成此目標需要建立一套整合性的逐步處理流程,從輸入資料的基本可靠性開始,逐漸推進至針對特定領域的模型適應策略。本論文透過一套端到端研究框架解決此挑戰,系統性地連結資料精煉與彈性的生成式模型適應方法。
    此流程的第一個關鍵階段著重於從資料來源提升資料品質。真實世界資料集,尤其是在醫療、商業分析以及特定領域文字分析等專業領域中的資料,經常存在缺失值與雜訊問題,進而降低後續人工智慧能力,例如 NER、RE 與 Text-to-SQL 任務的效能。為建立可靠的資料基礎,本論文提出一套由統計理論引導的缺失資料補值框架。此方法並非僅依賴資料驅動式最佳化,而是透過正式統計原理規範深度學習模型,使其在資料重建過程中能夠保留原始資料的統計分布特性。透過商業分析、醫療紀錄以及文字分類任務的大量實驗結果顯示,在初始階段建立高品質且具數學基礎的資料,能夠顯著提升後續預測準確度,確保後續模型適應程序建立於穩固且可靠的資料完整性基礎之上。
    在完成資料精煉基礎後,第二階段探討生成式模型與嵌入模型在基礎模型微調成本過高或受到限制時,如何適應特定使用者情境。為了有效連結原始使用者資料與預訓練基礎模型,本論文提出一套輔助模型訓練框架。不同於重新訓練或修改固定參數的生成式模型或大型語言模型,此方法透過部署輕量化輔助網路,作為智慧化中介層來完成知識轉移。藉由利用基礎模型與領域特定資料之間的共現關係以及潛在概念對齊,輔助模型能夠整合外部使用者知識與預訓練模型中的表示能力。此設計可在完整保留大型預訓練架構所蘊含的廣泛通用智慧的同時,實現高效率且低成本的領域適應。對於可以直接更新基礎模型參數的情境,本流程的最後階段進一步探討開放參數生成式模型在重疊領域中的微調挑戰。預訓練模型本身具有跨領域互相連結的知識結構,當其適應高度專業化的使用者資料時,往往會產生模糊的知識邊界。為了解決此問題,本論文提出一套結合軟式分群機制的遞迴學習框架。此策略透過反覆精煉方式,使廣泛的預訓練表示逐步朝向目標使用者領域進行聚焦,同時保留互補性的背景知識。藉由動態建模重疊概念,遞迴學習框架能夠實現平滑的語意轉換並避免災難性遺忘,使生成式模型能夠達成深度領域專業化。
    綜合而言,本論文證明提升生成式人工智慧對終端使用者的價值,需要採用整合性的完整生命週期方法,而非僅依靠計算規模的擴展。透過依序建立由統計補值方法所支撐的嚴謹資料品質、利用輔助對齊方法連結固定參數模型,以及透過遞迴式領域適應方法持續精煉開放參數模型,本研究提出一套全面性的架構,用於建構可靠、可適應且具領域感知能力的生成式人工智慧系統,使其能夠有效處理 NER、RE、情緒偵測以及 Text-to-SQL 等複雜任務。

    The dissertation aims to improve the knowledge inside the AI Model for different user data areas. The AI Model can be a deep learning model which was pre-trained on another dataset, such as a Large Language Model (LLM), a generative model, or a pre-trained model for embedding. After being improved through this framework, the upgraded model performs significantly better across specific downstream tasks, including Named Entity Recognition (NER), Relation Extraction (RE), Emotion Detection/Prediction, and Text-to-SQL translation tailored to specialized domain needs.
    As artificial intelligence transitions from general-purpose capabilities to specialized real-world deployment, a critical gap emerges expanding model scale alone does not guarantee utility for specialized applications. To deliver maximum benefit to end-users, AI systems must effectively bridge raw domain inputs with complex generative capabilities. Achieving this requires a unified, step-by-step pipeline, beginning with the fundamental reliability of the input data and progressing through targeted domain adaptation strategies. This dissertation addresses this challenge through an end-to-end research framework that systematically connects data refinement to flexible generative model adaptation.
    The first critical stage of this pipeline addresses data quality at the source. Real-world datasets, especially those in specialized fields like healthcare, business, and domain-specific text analysis, frequently suffer from missing values and noise, which corrupt downstream AI capabilities such as NER, RE, and Text-to-SQL. To construct a reliable foundation, this dissertation introduces a statistical theory-guided missing data imputation framework. By regulating deep learning models with formal statistical principles rather than relying purely on data-driven optimization, this approach preserves underlying statistical distributions during reconstruction. Extensive experiments across business analytics, medical records, and text classification demonstrate that establishing high-quality, mathematically sound data at the outset significantly enhances downstream predictive accuracy, ensuring that subsequent model adaptation relies on a solid baseline of data integrity.
    After the refined data foundation, the second stage examines how generative and embedding models can adapt to specific user contexts when fine-tuning the base model is either computationally prohibitive or restricted. To seamlessly connect raw user data with pre-trained foundation models, this dissertation introduces an Auxiliary Model Training framework. Rather than retraining or modifying the parameters of a frozen generative model or LLM, this method deploys a lightweight auxiliary network that acts as an intelligent intermediary. By leveraging co-occurrence relationships and latent concept alignment between the foundation model and domain-specific data, the auxiliary model harmonizes external user knowledge with pre-trained representations. This design enables efficient, low-overhead domain adaptation while fully preserving the broad general intelligence encoded in large, pre-trained architectures. For scenarios where base parameters can be directly updated, the final stage of the pipeline addresses the challenge of fine-tuning open-parameter generative models across overlapping domains. Pre-trained models inherently contain interconnected, cross-domain knowledge that often creates ambiguous boundary conditions when adapting to highly specialized user data. To resolve this, this dissertation proposes a Recursive Learning framework featuring a soft clustering mechanism. This strategy iteratively refines and sharpens broad pre-trained representations toward target user domains while preserving complementary background knowledge. By dynamically modeling overlapping concepts, the recursive framework enables smooth semantic transitions and prevents catastrophic forgetting, allowing generative models to achieve deep domain specialization.
    Collectively, this dissertation demonstrates that maximizing the end-user value of generative AI requires an integrated lifecycle approach rather than mere computational scaling. By sequentially establishing rigorous data quality through statistical imputation, bridging frozen models via auxiliary alignment, and iteratively refining open-parameter models through recursive domain adaptation, this research presents a comprehensive paradigm for building reliable, adaptable, and domain-aware generative AI systems that excel at complex tasks like NER, RE, Emotion Detection, and Text-to-SQL

    摘要 ii Abstract iv ACKNOWLEDGEMENTS vi TABLE OF CONTENTS vii LIST OF TABLES ix LIST OF FIGURES x CHAPTER 1 INTRODUCTION 1 1.1 Missing value problem in data collection 2 1.2 Language models for entity extraction and relation 4 1.3 Generative Learning methods for common supervised data 6 CHAPTER 2 MOTIVATION AND RESARCH OBJECTIVES 10 2.1 Motivation 10 2.2 Research Objectives 11 2.2.1 Foundational Data Refinement (Data Quality Stage) 11 2.2.2 Domain-Aware Model Adaptation (Model Fine-Tuning & Knowledge Alignment Stage) 11 2.2.3 Optimized End-to-End System Deployment (Output & Inference Stage) 11 2.3 Expected benefits and contributions 12 CHAPTER 3 LITERATURE REVIEW 13 3.1 Common approaches used in data quality improvement 13 3.2 Methods applied for NER and RE tasks. 14 3.3 Popular generative learning methods 15 3.3.1 Generative models for Emotion Detection 16 3.3.2 Generative models for Text-to-SQL 16 CHAPTER 4 TECHNICAL FRAMEWORK 18 4.1 Conceptual Overview of the research 18 4.2 K-Tails WGAN Imputation Method 21 4.2.1 Data selection 22 4.2.2 K-Tails WGAN construction 23 4.2.3 K-Tails Loss Function 25 4.3 The mixture of Auxiliary models for generative models adaptation 30 4.3.1 NER task with co-occurrence probability calculation and fulfillment approach 31 4.3.2 Entity Relation Extraction with auxiliary models 33 4.4 Recursive Learning methods applied for large generative models 36 4.4.1 Soft clustering module 37 4.4.2 Domain adaptation with Recursive Learning 39 CHAPTER 5 EXPERIMENT RESULTS 42 5.1 The effectiveness of K-Tails WGAN imputation 42 5.1.1 Dataset and method configurations 42 5.1.2 Metrics 43 5.1.3 Performance 44 5.2 Evaluation results and ablation studies of auxiliary methods for NER and RE tasks on the BioCreative Competition datasets 48 5.2.1 NER evaluation 48 5.2.2 RE evaluation 49 5.3 Experimental results and ablation studies of Recursive Learning 51 5.3.1 Emotion detection scenario 52 5.3.2 Text-to-SQL scenario 58 CHAPTER 6 LIMITATIONS AND FUTURE WORKS 64 CHAPTER 7 DISCUSSION AND CONCLUSION 66 REFERENCES 70

    Amati, G. (2009). BM25. In L. LIU & M. T. ÖZSU (Eds.), Encyclopedia of Database Systems (pp. 257–260). Springer US. https://doi.org/10.1007/978-0-387-39940-9_921
    Amberger, J. S., Bocchini, C. A., Schiettecatte, F., Scott, A. F., & Hamosh, A. (2015). OMIM.org: Online Mendelian Inheritance in Man (OMIM®), an Online catalog of human genes and genetic disorders. Nucleic Acids Research, 43(D1), D789–D798. https://doi.org/10.1093/nar/gku1205
    Amendolia, S. R., Cossu, G., Ganadu, M. L., Golosio, B., Masala, G. L., & Mura, G. M. (2003). A comparative study of K-Nearest Neighbour, Support Vector Machine and Multi-Layer Perceptron for Thalassemia screening. Chemometrics and Intelligent Laboratory Systems, 69(1–2). https://doi.org/10.1016/S0169-7439(03)00094-7
    Arjovsky, M., Chintala, S., & Bottou, L. (2017a). Wasserstein generative adversarial networks. 34th International Conference on Machine Learning, ICML 2017, 1.
    Arjovsky, M., Chintala, S., & Bottou, L. (2017b). Wasserstein generative adversarial networks. 34th International Conference on Machine Learning, ICML 2017, 1.
    Austin, P. C., White, I. R., Lee, D. S., & van Buuren, S. (2021). Missing Data in Clinical Research: A Tutorial on Multiple Imputation. In Canadian Journal of Cardiology (Vol. 37, Number 9). https://doi.org/10.1016/j.cjca.2020.11.010
    Bairoch, A. (2018). The cellosaurus, a cell-line knowledge resource. Journal of Biomolecular Techniques, 29(2). https://doi.org/10.7171/jbt.18-2902-002
    Bhatnagar, R., Sardar, S., Beheshti, M., & Podichetty, J. T. (2022). How can natural language processing help model informed drug development?: a review. In JAMIA Open (Vol. 5, Number 2). https://doi.org/10.1093/jamiaopen/ooac043
    Bianco, S., Celona, L., Donzella, M., & Napoletano, P. (2023). Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion. http://arxiv.org/abs/2306.11593
    Bohanec, M. (1988). Car Evaluation.
    Borg, A., Schiött, J., Ivegren, W., Gentline, C., Huss, V., Hugelius, A., Jobs, B., Espinosa, F., Ruiz, M., Edelbring, S., Georg, C., Skantze, G., & Parodis, I. (2025). AI-Enhanced Social Robotic Versus Computer-Based Virtual Patients for Clinical Reasoning Training in Medical Education: Observational Crossover Cohort Study. Journal of Medical Internet Research, 27. https://doi.org/10.2196/82541
    Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 2020-December.
    Caruccio, L., Cirillo, S., Polese, G., Solimando, G., Sundaramurthy, S., & Tortora, G. (2024). Claude 2.0 large language model: Tackling a real-world classification problem with a new iterative prompt engineering approach. Intelligent Systems with Applications, 21. https://doi.org/10.1016/j.iswa.2024.200336
    Chang, S., & Fosler-Lussier, E. (2023). How to Prompt LLMs for Text-to-SQL: A Study in Zero-shot, Single-domain, and Cross-domain Settings. NeurIPS 2023. https://doi.org/https://doi.org/10.48550/arXiv.2305.11853
    Curnow, E., Cornish, R. P., Heron, J. E., Carpenter, J. R., & Tilling, K. (2024). Multiple imputation using auxiliary imputation variables that only predict missingness can increase bias due to data missing not at random. BMC Medical Research Methodology, 24(1). https://doi.org/10.1186/s12874-024-02353-9
    Deforth, M., Heinze, G., & Held, U. (2024). The performance of prognostic models depended on the choice of missing value imputation algorithm: a simulation study. Journal of Clinical Epidemiology, 176. https://doi.org/10.1016/j.jclinepi.2024.111539
    Demirtas, H. (2018). Flexible Imputation of Missing Data. Journal of Statistical Software, 85(Book Review 4). https://doi.org/10.18637/jss.v085.b04
    Demner-Fushman, D., Chapman, W. W., & McDonald, C. J. (2009). What can natural language processing do for clinical decision support? In Journal of Biomedical Informatics (Vol. 42, Number 5). https://doi.org/10.1016/j.jbi.2009.08.007
    Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, 1.
    Do, D. T., Yang, M. R., Lam, L. H. T., Le, N. Q. K., & Wu, Y. W. (2022). Improving MGMT methylation status prediction of glioblastoma through optimizing radiomics features using genetic algorithm-based machine learning approach. Scientific Reports, 12(1). https://doi.org/10.1038/s41598-022-17707-w
    Dong, W., Fong, D. Y. T., Yoon, J. sun, Wan, E. Y. F., Bedford, L. E., Tang, E. H. M., & Lam, C. L. K. (2021). Generative adversarial networks for imputing missing data for big data clinical research. BMC Medical Research Methodology, 21(1). https://doi.org/10.1186/s12874-021-01272-3
    Dzabraev, M., Kunitsyn, A., & Ivaniuta, A. (2024). VLRM: Vision-Language Models act as Reward Models for Image Captioning. http://arxiv.org/abs/2404.01911
    Finsel, J. S., Axelrad, H., Choi, S. J., Derous, E., Gu, X., Guandalini, P., Ha, J., Kim, E. S., Marzec, I., Mykletun, R. J., Oliveira, E., Pajic, S., Schellaert, M., Van der Heijden, B. I. J. M., Vignoli, M., Wöhrmann, A. M., & Deller, J. (2025). Development and validation of the short form of the Later Life Workplace Index: a study across 10 countries. Work, Aging and Retirement. https://doi.org/10.1093/workar/waaf006
    Gao, D., Wang, H., Li, Y., Sun, X., Qian, Y., Ding, B., & Zhou, J. (2024). Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation. Proc. VLDB Endow., 17(5), 1132–1145. https://doi.org/10.14778/3641204.3641221
    Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. Advances in Neural Information Processing Systems, 3(January), 2672–2680. https://doi.org/10.1007/978-3-658-40442-0_9
    Google. (2024). Gemini Sentiment Analysis Gallery. Google AI Studio.
    Graham, J. W. (2009). Missing data analysis: Making it work in the real world. In Annual Review of Psychology (Vol. 60). https://doi.org/10.1146/annurev.psych.58.110405.085530
    Guu, K., Lee, K., Tung, Z., Pasupat, P., & Chang, M. W. (2020). REALM: Retrieval-Augmented language model pre-training. 37th International Conference on Machine Learning, ICML 2020, PartF168147-6.
    Han, J., Gong, K., Zhang, Y., Wang, J., Zhang, K., Lin, D., Qiao, Y., Gao, P., & Yue, X. (n.d.). OneLLM: One Framework to Align All Modalities with Language. Retrieved https://github.com/csuhan/OneLLM
    Han, P., Li, X., Wang, X., Wang, S., Gao, C., & Chen, W. (2022). Exploring the effects of drug, disease, and protein dependencies on biomedical named entity recognition: A comparative analysis. Frontiers in Pharmacology, 13. https://doi.org/10.3389/fphar.2022.1020759
    Herrero-Zazo, M., Segura-Bedmar, I., Martínez, P., & Declerck, T. (2013). The DDI corpus: An annotated corpus with pharmacological substances and drug-drug interactions. Journal of Biomedical Informatics, 46(5). https://doi.org/10.1016/j.jbi.2013.07.011
    Hong, Z., Yuan, Z., Zhang, Q., Chen, H., Dong, J., Huang, F., & Huang, X. (2024). Next-Generation Database Interfaces: A Survey of LLM-based Text-to-SQL. http://arxiv.org/abs/2406.08426
    Hopkins, M., Reeber, E., Forman, G., & Suermondt, J. (1999). Spambase data set. Hewlett-Packard Labs, 41(8).
    Hu, W., Xu, Y., Li, Y., Li, W., Chen, Z., & Tu, Z. (2024). BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions. www.aaai.org
    Hutcheson, G. (2012). Missing Data: data replacement and imputation. Journal of Modelling in Management, 7(2). https://doi.org/10.1108/jm2.2012.29707baa.002
    Islamaj, R., Lai, P.-T., Wei, C.-H., Luo, L., & Lu, Z. (2023). The overview of the BioRED (Biomedical Relation Extraction Dataset) track at BioCreative VIII. Zenodo. https://doi.org/10.5281/zenodo.10351131
    Izacard, G., & Grave, E. (2021). Leveraging passage retrieval with generative models for open domain question answering. EACL 2021 - 16th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the Conference. https://doi.org/10.18653/v1/2021.eacl-main.74
    Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., Dwivedi-Yu, J., Joulin, A., Riedel, S., & Grave, E. (2023). Atlas: few-shot learning with retrieval augmented language models. J. Mach. Learn. Res., 24(1).
    Jin, P., Takanobu, R., Zhang, W., Cao, X., & Yuan, L. (2024). Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13700–13710. https://doi.org/10.1109/CVPR52733.2024.01300
    Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W. T. (2020). Dense passage retrieval for open-domain question answering. EMNLP 2020 - 2020 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference. https://doi.org/10.18653/v1/2020.emnlp-main.550
    Khan, S. I., & Hoque, A. S. M. L. (2020). SICE: an improved missing data imputation technique. Journal of Big Data, 7(1). https://doi.org/10.1186/s40537-020-00313-w
    Kim, J. D., Ohta, T., Tateisi, Y., & Tsujii, J. (2003). GENIA corpus - A semantically annotated corpus for bio-textmining. Bioinformatics, 19(SUPPL. 1). https://doi.org/10.1093/bioinformatics/btg1023
    Kong, W., Wong, B. J. H., Hui, H. W. H., Lim, K. P., Wang, Y., Wong, L., & Goh, W. W. Bin. (2023). ProJect: a powerful mixed-model missing value imputation method. Briefings in Bioinformatics, 24(4). https://doi.org/10.1093/bib/bbad233
    Koren, Y., Bell, R., & Volinsky, C. (2009). Matrix factorization techniques for recommender systems. Computer, 42(8). https://doi.org/10.1109/MC.2009.263
    Krallinger, M., Leitner, F., Rodriguez-Penagos, C., & Valencia, A. (2008). Overview of the protein-protein interaction annotation extraction task of BioCreative II. In Genome Biology (Vol. 9, Number SUPPL. 2). https://doi.org/10.1186/gb-2008-9-s2-s4
    Lai, P. T., & Lu, Z. (2020). BERT-GT: Cross-sentence n-ary relation extraction with BERT and Graph Transformer. Bioinformatics, 36(24). https://doi.org/10.1093/bioinformatics/btaa1087
    Lai, P.-T., Wei, C.-H., Luo, L., Chen, Q., & Lu, Z. (2023). BioREx: Improving biomedical relation extraction by leveraging heterogeneous datasets. Journal of Biomedical Informatics, 146. https://doi.org/10.1016/j.jbi.2023.104487
    Lee, C. H., Polozov, O., & Richardson, M. (2021). KaggleDBQA: Realistic evaluation of text-to-SQL parsers. ACL-IJCNLP 2021 - 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Proceedings of the Conference, 1. https://doi.org/10.18653/v1/2021.acl-long.176
    Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4). https://doi.org/10.1093/bioinformatics/btz682
    Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W. T., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 2020-December.
    Li, H., Zhang, J., Li, C., & Chen, H. (2023). RESDSQL: Decoupling Schema Linking and Skeleton Parsing for Text-to-SQL. Proceedings of the 37th AAAI Conference on Artificial Intelligence, AAAI 2023, 37. https://doi.org/10.1609/aaai.v37i11.26535
    Li, H., Zhang, J., Liu, H., Fan, J., Zhang, X., Zhu, J., Wei, R., Pan, H., Li, C., & Chen, H. (2024). CodeS: Towards Building Open-source Language Models for Text-to-SQL. Proc. ACM Manag. Data, 2(3). https://doi.org/10.1145/3654930
    Li, J., Hui, B., Qu, G., Yang, J., Li, B., Li, B., Wang, B., Qin, B., Geng, R., Huo, N., Zhou, X., Ma, C., Li, G., Chang, K. C. C., Huang, F., Cheng, R., & Li, Y. (2023). Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs. Advances in Neural Information Processing Systems, 36.
    Li, J., Sun, Y., Johnson, R. J., Sciaky, D., Wei, C. H., Leaman, R., Davis, A. P., Mattingly, C. J., Wiegers, T. C., & Lu, Z. (2016). BioCreative V CDR task corpus: a resource for chemical disease relation extraction. Database : The Journal of Biological Databases and Curation, 2016. https://doi.org/10.1093/database/baw068
    Lichman, M. (2013). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. In UCI Machine Learning Repository.
    Liu, Z., Yang, K., Xie, Q., Zhang, T., & Ananiadou, S. (2024). EmoLLMs: A Series of Emotional Large Language Models and Annotation Tools for Comprehensive Affective Analysis. Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’24, 5487–5496. https://doi.org/10.1145/3637528.3671552
    Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., & Roberts, A. (2023). The Flan Collection: Designing Data and Methods for Effective Instruction Tuning. Proceedings of Machine Learning Research, 202.
    Luo, L., Lai, P. T., Wei, C. H., Arighi, C. N., & Lu, Z. (2022). BioRED: A rich biomedical relation extraction dataset. In Briefings in Bioinformatics (Vol. 23, Number 5). https://doi.org/10.1093/bib/bbac282
    Luo, L., Wei, C. H., Lai, P. T., Leaman, R., Chen, Q., & Lu, Z. (2023). AIONER: all-in-one scheme-based biomedical named entity recognition using deep learning. Bioinformatics, 39(5). https://doi.org/10.1093/bioinformatics/btad310
    Lyu, C., Wu, M., Wang, L., Huang, X., Liu, B., Du, Z., Shi, S., & Tu, Z. (2023). Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration. http://arxiv.org/abs/2306.09093
    Masoudnia, S., & Ebrahimpour, R. (2014). Mixture of experts: A literature survey. Artificial Intelligence Review, 42(2). https://doi.org/10.1007/s10462-012-9338-y
    Min, S., Lewis, M., Zettlemoyer, L., & Hajishirzi, H. (2022). MetaICL: Learning to Learn In Context. NAACL 2022 - 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference. https://doi.org/10.18653/v1/2022.naacl-main.201
    Neumann, M., King, D., Beltagy, I., & Ammar, W. (2019). ScispaCy: Fast and robust models for biomedical natural language processing. BioNLP 2019 - SIGBioMed Workshop on Biomedical Natural Language Processing, Proceedings of the 18th BioNLP Workshop and Shared Task. https://doi.org/10.18653/v1/w19-5034
    OpenAI. (2023). OpenAI - GPT-4. Https://Openai.Com/Gpt-4.
    Öztürk, H., Özgür, A., Schwaller, P., Laino, T., & Ozkirimli, E. (2020). Exploring chemical space using natural language processing methodologies for drug discovery. In Drug Discovery Today (Vol. 25, Number 4). https://doi.org/10.1016/j.drudis.2020.01.020
    Phan, C.-P., Phan, B., & Chiang, J.-H. (2024). Optimized biomedical entity relation extraction method with data augmentation and classification using GPT-4 and Gemini. Database, 2024, baae104. https://doi.org/10.1093/database/baae104
    Potthoff, R. F., Tudor, G. E., Pieper, K. S., & Hasselblad, V. (2006). Can one assess whether missing data are missing at random in medical studies? Statistical Methods in Medical Research, 15(3). https://doi.org/10.1191/0962280206sm448oa
    Pourreza, M., & Rafiei, D. (2023). DIN-SQL: decomposed in-context learning of text-to-SQL with self-correction. Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23.
    Pyysalo, S., Ginter, F., Heimonen, J., Björne, J., Boberg, J., Järvinen, J., & Salakoski, T. (2007). BioInfer: A corpus for information extraction in the biomedical domain. BMC Bioinformatics, 8. https://doi.org/10.1186/1471-2105-8-50
    Qu, G., Li, J., Li, B., Qin, B., Huo, N., Ma, C., & Cheng, R. (n.d.). Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation.
    Radford, A., Wook, J., Chris, K., Aditya, H., Gabriel, R., Sandhini, G., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2019). CLIP: Learning Transferable Visual Models From Natural Language Supervision. In OpenAI.
    Rahman, M. M., & Davis, D. N. (2013). Machine learning-based missing value imputation method for clinical datasets. Lecture Notes in Electrical Engineering, 229 LNEE. https://doi.org/10.1007/978-94-007-6190-2_19
    Rajkumar, N., Li, R., & Bahdanau, D. (2022). Evaluating the Text-to-SQL Capabilities of Large Language Models. http://arxiv.org/abs/2204.00498
    Reyes-Ortiz, J. A., Gonzalez-Beltran, B. A., & Gallardo-Lopez, L. (2016). Clinical Decision Support Systems: A Survey of NLP-Based Approaches from Unstructured Data. Proceedings - International Workshop on Database and Expert Systems Applications, DEXA, 2016-Febru. https://doi.org/10.1109/DEXA.2015.47
    Rubin, D. B. (1986). Statistical matching using file concatenation with adjusted weights and multiple imputations. Journal of Business and Economic Statistics, 4(1). https://doi.org/10.1080/07350015.1986.10509497
    Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Le Scao, T., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N. V., … Rush, A. M. (2022). Multitask prompted training enables zero-shot task generalization. ICLR 2022 - 10th International Conference on Learning Representations.
    Schoch, C. L., Ciufo, S., Domrachev, M., Hotton, C. L., Kannan, S., Khovanskaya, R., Leipe, D., McVeigh, R., O’Neill, K., Robbertse, B., Sharma, S., Soussov, V., Sullivan, J. P., Sun, L., Turner, S., & Karsch-Mizrachi, I. (2020). NCBI Taxonomy: A comprehensive update on curation, resources and tools. In Database (Vol. 2020). https://doi.org/10.1093/database/baaa062
    Scholak, T., Schucher, N., & Bahdanau, D. (2021). PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from Language Models. EMNLP 2021 - 2021 Conference on Empirical Methods in Natural Language Processing, Proceedings. https://doi.org/10.18653/v1/2021.emnlp-main.779
    Shah, A. D., Bartlett, J. W., Carpenter, J., Nicholas, O., & Hemingway, H. (2014). Comparison of random forest and parametric imputation models for imputing missing data using MICE: A CALIBER study. American Journal of Epidemiology, 179(6). https://doi.org/10.1093/aje/kwt312
    Shamsuddeen Hassan Muhammad, Seid Muhie Yimam, Nedjma OUSIDHOUM, Idris Abdulmumin, Ibrahim Said Ahmad, Alham Fikri Aji, David Ifeoluwa Adelani, Vladimir Araujo, Abinew Ali Ayele, Tadesse Destaw Belay, Daniela Teodorescu, Jan Philip Wahle, Terry Ruas, Nirmal Surange, & Yi Zhou. (2025). ACL SemEval 2025 - Task 11. ACL SemEval 2025.
    Sherry, S. T., Ward, M. H., Kholodov, M., Baker, J., Phan, L., Smigielski, E. M., & Sirotkin, K. (2001). DbSNP: The NCBI database of genetic variation. Nucleic Acids Research, 29(1). https://doi.org/10.1093/nar/29.1.308
    Stekhoven, D. J., & Bühlmann, P. (2012). Missforest-Non-parametric missing value imputation for mixed-type data. Bioinformatics, 28(1). https://doi.org/10.1093/bioinformatics/btr597
    Sterner, I., Lin, W., Chen, J., & Byrne, B. (2024). Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches. http://arxiv.org/abs/2403.11317
    Talaei, S., Pourreza, M., Chang, Y.-C., Mirhoseini, A., & Saberi, A. (2024). CHESS: Contextual Harnessing for Efficient SQL Synthesis. http://arxiv.org/abs/2405.16755
    TIANCHI. (2023). Age Assessment & Disease Risk Prediction H5. https://www.kaggle.com/datasets/marquis03/age-assessment-and-disease-risk-prediction-h5/discussion/456430
    van Buuren, S., & Groothuis-Oudshoorn, K. (2011). mice: Multivariate imputation by chained equations in R. Journal of Statistical Software, 45(3). https://doi.org/10.18637/jss.v045.i03
    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 2017-Decem.
    Wallace, M. L., Anderson, S. J., & Mazumdar, S. (2010). A stochastic multiple imputation algorithm for missing covariate data in tree-structured survival analysis. Statistics in Medicine, 29(29). https://doi.org/10.1002/sim.4079
    Wang, B., Shin, R., Liu, X., Polozov, O., & Richardson, M. (2020). RAT-SQL: Relation-aware schema encoding and linking for text-to-SQL parsers. Proceedings of the Annual Meeting of the Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.677
    Wei, C. H., Allot, A., Leaman, R., & Lu, Z. (2019). PubTator central: automated concept annotation for biomedical full text articles. Nucleic Acids Research, 47(W1). https://doi.org/10.1093/nar/gkz389
    Wikimedia Foundation. (2025). MediaWiki Action API. https://www.mediawiki.org/wiki/API:Main_page
    Wu, H., Zhang, Z., Zhang, E., Chen, C., Liao, L., Wang, A., Xu, K., Li, C., Hou, J., Zhai, G., Xue, G., Sun, W., Yan, Q., & Lin, W. (2024). Q-Instruct: Improving Low-Level Visual Abilities for Multi-Modality Foundation Models. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 25490–25500. https://doi.org/10.1109/CVPR52733.2024.02408
    Yoon, J., Jordon, J., & Van Der Schaar, M. (2018). GAIN: Missing data imputation using generative adversarial nets. 35th International Conference on Machine Learning, ICML 2018, 13.
    Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., Zhang, Z., & Radev, D. R. (2018). Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018. https://doi.org/10.18653/v1/d18-1425
    Zhang, H., Cao, R., Chen, L., Xu, H., & Yu, K. (n.d.). ACT-SQL: In-Context Learning for Text-to-SQL with Automatically-Generated Chain-of-Thought. Retrieved https://platform.openai.com/examples/
    Zhang, S. (2012). Nearest neighbor selection for iteratively kNN imputation. Journal of Systems and Software, 85(11). https://doi.org/10.1016/j.jss.2012.05.073
    Ziletti, A., & DAmbrosi, L. (2024). Retrieval augmented text-to-SQL generation for epidemiological question answering using electronic health records. In T. Naumann, A. Ben Abacha, S. Bethard, K. Roberts, & D. Bitterman (Eds.), Proceedings of the 6th Clinical Natural Language Processing Workshop (pp. 47–53). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.clinicalnlp-1.4
    Zwitter, M., & Soklic, M. (2009). UCI Machine Learning Repository: Breast Cancer Data Set. In https://archive.ics.uci.edu/ml/datasets/Breast+Cancer.

    QR CODE