| 研究生: |
王宇軒 Wang, Yu-Hsuan |
|---|---|
| 論文名稱: |
基於迭代式模式歸納和語意實體解析之圖檢索增強生成框架:以半結構化資料為例 A Graph Retrieval-Augmented Generation Framework Enhanced by Iterative Schema Induction and Semantic Entity Resolution for Semi-Structured Data |
| 指導教授: |
蔣榮先
Chiang, Jung-Hsien |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 醫學資訊研究所 Institute of Medical Informatics |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 110 |
| 中文關鍵詞: | 大型語言模型 、檢索增強生成 、圖檢索增強生成 、知識圖譜 、半結構化資料 、知識抽取 、模式歸納 、實體解析 、多文件問答 |
| 外文關鍵詞: | Large Language Models, Retrieval-Augmented Generation, Graph-based Retrieval-Augmented Generation, Knowledge Graph, Semi-structured Data, Knowledge Extraction, Schema Induction, Entity Resolution, Multi-document QA |
| 相關次數: | 點閱:2 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
在現實醫療與企業場域中,持續累積的營運紀錄多屬於半結構化資料。此類資料以實體與時序為單位記錄,使得同一實體的相關資訊往往散落於多筆、由不同人員書寫的條目之中。因此,回答需要跨越多筆紀錄聚合證據的問題,成為這類資料上有價值卻困難的任務。傳統檢索增強生成系統在此類跨紀錄聚合問題上表現受限,其根本原因在於將文本塊視為彼此獨立的單位,缺乏顯式表徵跨紀錄實體關聯的機制,難以支援需要彙整多筆證據的查詢。圖檢索增強生成框架則以實體與關係為核心,將原本隱含於非結構化文本、分散於不同紀錄的同一實體關聯經由圖譜重組而顯性化,把散落的紀錄連結為可遍歷的結構,因而較適合處理跨紀錄聚合問題。然而,圖結構所提供的這項表徵能力,是否能轉化為實際的下游效益,取決於能否從雜亂的原始資料建構出高品質的知識圖譜。
知識抽取與實體解析為圖譜建構階段的核心步驟,目的分別在於將原始資料轉化為圖譜結構,以及對初始圖譜進行實體層次的整併。現有知識抽取多仰賴大型語言模型抽取實體與關係,而在無模式、固定模式與專家制定模式之間,各有圖譜品質或人力成本的取捨。考量到人力成本與效率以及領域模式的適應性,現實場景中需要一套能根據領域資料自動歸納模式的流程。傳統實體解析做法多流於合併完全相同的實體,或透過預設規則進行特殊符號處理,在面對真實世界的雜亂資料時,紀錄書寫方式的不一致常產生不同字面表述的實體,使得同一實體在圖中被切分為多個互不相連的節點,進而削弱跨紀錄的證據聚合,此類字面變異難以預先以規則窮舉,需在語意層次加以解析。
為克服上述限制,本研究在圖檢索增強生成框架下,提出一套適用於真實半結構化資料的圖譜建構方法,包含兩個建構階段模組。「迭代式模式歸納」在零樣本且免人工標註的前提下,自動依領域資料歸納出抽取模式,進而引導大型語言模型完成下游知識抽取與初始圖譜建構。「語意實體解析」藉由實體名稱與描述的語意相似度,自動合併字面相異的等價實體。為使這兩個模組的效果能被清楚地隔離與衡量,本研究選用建構流程相對基礎的圖檢索增強生成框架作為實驗載體,在固定檢索與生成設定的條件下,單獨檢驗建構階段改動對下游問答的貢獻。
為驗證本框架回答跨紀錄聚合問題的能力,本研究蒐集真實醫療場域的護理紀錄與企業技術服務工單作為半結構化資料,充當檢索增強生成框架的外部知識庫與問答集資料來源。主要實驗比較三種層級:不使用圖結構、使用圖結構、以及優化建構品質後的圖結構。結果顯示,圖結構的加入使下游任務在兩個資料集上平均提升約三個百分點,更進一步的圖譜建構優化可以在下游任務上再帶來平均約三個百分點的提升,證明建構品質確實會影響下游表現。本研究進一步以消融實驗分析兩模組的個別與疊加貢獻:迭代式模式歸納彌補了無模式、固定預設模式、專家制定模式在跨領域使用上的限制;語意實體解析則將被字面變異切分的節點重新整併,確保跨紀錄實體資訊的完整性。實驗證明強化圖譜建構階段的設計能提升下游任務表現,本框架的優勢主要集中於跨紀錄聚合的情境。
綜上,本研究在圖檢索增強生成框架下,針對真實半結構化資料在圖譜建構階段的限制提出改善方法,透過跨紀錄聚合問題的實驗進行回答能力的驗證,並以消融實驗證明了建構階段優化的有效性。本框架對持續累積、紀錄間關聯隱含的半結構化資料,於跨紀錄知識檢索問答上具實際應用潛力。
In real-world healthcare and enterprise domains, continuously accumulated operational records are often semi-structured. Such data are recorded around entities and temporal events, so that information about the same entity is frequently scattered across multiple entries written by different personnel. As a result, answering questions that require aggregating evidence across multiple records becomes a valuable yet challenging task on such data. Traditional retrieval-augmented generation systems are limited in such multi-document QA problems, fundamentally because they treat text chunks as mutually independent units and lack an explicit mechanism for representing cross-record entity relationships, making it difficult to support queries that require synthesizing evidence from multiple records. Graph-based retrieval-augmented generation frameworks, by contrast, are centered on entities and relations. By reconstructing into a graph the associations of the same entity that are originally implicit in unstructured text and dispersed across different records, they make these associations explicit and link scattered records into a traversable structure, and are therefore better suited to multi-document QA problems. However, whether the representational capability provided by the graph structure can be translated into actual downstream benefits depends on whether a high-quality knowledge graph can be constructed from noisy raw data.
Knowledge extraction and entity resolution are the core steps of the graph construction stage, aiming respectively to transform raw data into a graph structure and to perform entity-level consolidation of the initial graph. Existing knowledge extraction methods mostly rely on large language models to extract entities and relations, and involve trade-offs between graph quality and labor cost among schema-free, fixed-schema, and expert-defined-schema approaches. Considering labor cost and efficiency as well as the adaptability of domain schemas, practical scenarios call for a process that can automatically induce a schema from domain data. Conventional entity resolution methods, in turn, tend to merge only exactly identical entities or to handle special symbols through predefined rules. When facing noisy real-world data, inconsistent writing styles frequently produce entities with different surface forms, causing the same entity to be split into multiple disconnected nodes in the graph and thereby weakening cross-record evidence aggregation; such surface-form variation is difficult to exhaustively enumerate with predefined rules and must instead be resolved at the semantic level.
To overcome the above limitations, this study proposes, under the graph-based retrieval-augmented generation framework, a graph construction method suited to real-world semi-structured data, comprising two construction-stage modules. Iterative Schema Induction, under a zero-shot and annotation-free setting, automatically induces an extraction schema from domain data and thereby guides the large language model in downstream knowledge extraction and initial graph construction. Semantic Entity Resolution automatically merges equivalent entities with different surface forms based on the semantic similarity of entity names and descriptions. To clearly isolate and measure the effects of these two modules, this study adopts a graph-based retrieval-augmented generation framework with a relatively basic construction pipeline as the experimental carrier, and, under fixed retrieval and generation settings, evaluates in isolation the contribution of construction-stage modifications to downstream question answering.
To validate the proposed framework's ability to answer multi-document QA problems, this study collects nursing records from a real healthcare setting and enterprise technical-service tickets as semi-structured data, serving as both the external knowledge base of the retrieval-augmented generation framework and the source of the question-answering set. The main experiments compare three levels: without the graph structure, with the graph structure, and with the graph structure after construction-quality optimization. The results show that introducing the graph structure improves downstream task performance by an average of about 3 percentage points across the two datasets, and further graph-construction optimization yields an additional average improvement of about 3 percentage points in downstream tasks, demonstrating that construction quality indeed affects downstream performance. This study further uses ablation experiments to analyze the individual and combined contributions of the two modules: Iterative Schema Induction compensates for the limitations of schema-free, fixed predefined, and expert-defined schemas in cross-domain use, while Semantic Entity Resolution reconsolidates nodes that were split by surface-form variation, ensuring the completeness of cross-record entity information. The experiments demonstrate that strengthening the design of the graph construction stage improves downstream task performance, and the advantage of the proposed framework is mainly concentrated in multi-document QA scenarios.
In summary, under the graph-based retrieval-augmented generation framework, this study proposes improvements addressing the limitations of the graph construction stage on real-world semi-structured data, validates its answering ability through experiments on multi-document QA problems, and verifies the effectiveness of construction-stage optimization through ablation experiments. The proposed framework has practical application potential for cross-record knowledge retrieval and question answering over continuously accumulated semi-structured data whose inter-record associations are implicit.
[1] Peter Buneman. Semistructured data. In Proceedings of the sixteenth ACM SIGACT-SIGMOD-SIGART symposium on Principles of database systems, PODS '97, pages 117--121, New York, NY, USA, 1997. Association for Computing Machinery.
[2] Serge Abiteboul. Querying Semi-Structured Data. In Proceedings of the 6th International Conference on Database Theory, ICDT '97, pages 1--18, Berlin, Heidelberg, 1997. Springer-Verlag.
[3] Pinjia He, Jieming Zhu, Zibin Zheng, and Michael R. Lyu. Drain: An Online Log Parsing Approach with Fixed Depth Tree. In 2017 IEEE International Conference on Web Services (ICWS), pages 33--40, June 2017.
[4] Joshua Mandel, David Kreda, Kenneth Mandl, Isaac Kohane, and Rachel Ramoni. SMART on FHIR: A standards-based, interoperable apps platform for electronic health records. Journal of the American Medical Informatics Association, 23:ocv189, February 2016.
[5] Teng Lin, Yuyu Luo, Honglin Zhang, Jicheng Zhang, Chunlin Liu, Kaishun Wu, and Nan Tang. MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering, September 2025. arXiv:2502.18993 [cs.CL].
[6] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen-tau Yih, Tim Rocktaschel, Sebastian Riedel, and Douwe Kiela. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, April 2021. arXiv:2005.11401 [cs].
[7] Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. From Local to Global: A Graph RAG Approach to Query-Focused Summarization, February 2025. arXiv:2404.16130 [cs].
[8] Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. Retrieval-Augmented Generation for Large Language Models: A Survey, March 2024. arXiv:2312.10997 [cs.CL].
[9] Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions, June 2023. arXiv:2212.10509 [cs.CL].
[10] Zirui Guo, Lianghao Xia, Yanhua Yu, Tu Ao, and Chao Huang. LightRAG: Simple and Fast Retrieval-Augmented Generation, April 2025. arXiv:2410.05779 [cs].
[11] Mohammad Sadeq Abolhasani, Yang Ba, Yixuan He, and Rong Pan. Beyond Predefined Schemas: TRACE-KG for Context-Enriched Knowledge Graphs from Complex Documents, April 2026. arXiv:2604.03496 [cs.AI].
[12] Jiaxin Bai, Wei Fan, Qi Hu, Qing Zong, Chunyang Li, Hong Ting Tsang, Hongyu Luo, Yauwai Yim, Haoyu Huang, Xiao Zhou, Feng Qin, Tianshi Zheng, Xi Peng, Xin Yao, Huiwen Yang, Leijie Wu, Yi Ji, Gong Zhang, Renhai Chen, and Yangqiu Song. AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora, May 2025.
[13] Haoyu Han, Li Ma, Yu Wang, Harry Shomer, Yongjia Lei, Zhisheng Qi, Kai Guo, Zhigang Hua, Bo Long, Hui Liu, Charu C. Aggarwal, and Jiliang Tang. RAG vs. GraphRAG: A Systematic Evaluation and Key Insights, March 2026. arXiv:2502.11371 [cs.IR].
[14] Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A. Smith, and Mike Lewis. Measuring and Narrowing the Compositionality Gap in Language Models, October 2023. arXiv:2210.03350 [cs.CL].
[15] Boci Peng, Yun Zhu, Yongchao Liu, Xiaohe Bo, Haizhou Shi, Chuntao Hong, Yan Zhang, and Siliang Tang. Graph Retrieval-Augmented Generation: A Survey, September 2024. arXiv:2408.08921 [cs.AI].
[16] Reinhard Moratz, Niklas Daute, James Ondieki, Markus Kattenbeck, Mario Krajina, and Ioannis Giannopoulos. Bilateral Spatial Reasoning about Street Networks: Graph-based RAG with Qualitative Spatial Representations, December 2025. arXiv:2512.15388 [cs.AI].
[17] Haoyu Han, Yu Wang, Harry Shomer, Kai Guo, Jiayuan Ding, Yongjia Lei, Mahantesh Halappanavar, Ryan A. Rossi, Subhabrata Mukherjee, Xianfeng Tang, Qi He, Zhigang Hua, Bo Long, Tong Zhao, Neil Shah, Amin Javari, Yinglong Xia, and Jiliang Tang. Retrieval-Augmented Generation with Graphs (GraphRAG), January 2025. arXiv:2501.00309 [cs.IR].
[18] Boyu Chen, Zirui Guo, Zidan Yang, Yuluo Chen, Junze Chen, Zhenghao Liu, Chuan Shi, and Cheng Yang. PathRAG: Pruning Graph-based Retrieval Augmented Generation with Relational Paths, November 2025. arXiv:2502.14902 [cs.CL].
[19] Qinggang Zhang, Shengyuan Chen, Yuanchen Bei, Zheng Yuan, Huachi Zhou, Zijin Hong, Hao Chen, Yilin Xiao, Chuang Zhou, Junnan Dong, Yi Chang, and Xiao Huang. A Survey of Graph Retrieval-Augmented Generation for Customized Large Language Models, September 2025. arXiv:2501.13958 [cs.CL].
[20] Samuel Thio, Matthew Lewis, Spiros Denaxas, and Richard JB Dobson. Unlocking Electronic Health Records: A Hybrid Graph RAG Approach to Safe Clinical AI for Patient QA. Frontiers in Digital Health, 8:1780700, March 2026. arXiv:2602.00009 [cs.CL].
[21] Amy J. Starmer, Nancy D. Spector, Rajendu Srivastava, Daniel C. West, Glenn Rosenbluth, April D. Allen, Elizabeth L. Noble, Lisa L. Tse, Anuj K. Dalal, Carol A. Keohane, Stuart R. Lipsitz, Jeffrey M. Rothschild, Matthew F. Wien, Catherine S. Yoon, Katherine R. Zigmont, Karen M. Wilson, Jennifer K. O'Toole, Lauren G. Solan, Megan Aylor, Zia Bismilla, Maitreya Coffey, Sanjay Mahant, Rebecca L. Blankenburg, Lauren A. Destino, Jennifer L. Everhart, Shilpa J. Patel, James F. Bale, Jaime B. Spackman, Adam T. Stevenson, Sharon Calaman, F. Sessions Cole, Dorene F. Balmer, Jennifer H. Hepps, Joseph O. Lopreiato, Clifton E. Yu, Theodore C. Sectish, and Christopher P. Landrigan. Changes in Medical Errors after Implementation of a Handoff Program. New England Journal of Medicine, 371(19):1803--1812, November 2014. Publisher: Massachusetts Medical Society _eprint: https://www.nejm.org/doi/pdf/10.1056/NEJMsa1405556.
[22] Adam Rule, Steven Bedrick, Michael F. Chiang, and Michelle R. Hribar. Length and Redundancy of Outpatient Progress Notes Across a Decade at an Academic Medical Center. JAMA Network Open, 4(7):e2115334, July 2021.
[23] Jan Clusmann, Fiona R. Kolbinger, Hannah Sophie Muti, Zunamys I. Carrero, Jan-Niklas Eckardt, Narmin Ghaffari Laleh, Chiara Maria Lavinia Loffler, Sophie-Caroline Schwarzkopf, Michaela Unger, Gregory P. Veldhuizen, Sophia J. Wagner, and Jakob Nikolas Kather. The future landscape of large language models in medicine. Communications Medicine, 3(1):141, October 2023. Publisher: Nature Publishing Group.
[24] Mohammad Baqar. RAG4Tickets: AI-Powered Ticket Resolution via Retrieval-Augmented Generation on JIRA and GitHub Data, October 2025.
[25] Yuxuan Chen, Dewen Guo, Sen Mei, Xinze Li, Hao Chen, Yishan Li, Yixuan Wang, Chaoyue Tang, Ruobing Wang, Dingjun Wu, Yukun Yan, Zhenghao Liu, Shi Yu, Zhiyuan Liu, and Maosong Sun. UltraRAG: A Modular and Automated Toolkit for Adaptive Retrieval-Augmented Generation, March 2025.
[26] Michael Wornow, Rahul Thapa, Ethan Steinberg, Jason A. Fries, and Nigam H. Shah. EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models, December 2023. arXiv:2307.02028 [cs.LG].
[27] Bhaskarjit Sarmah, Benika Hall, Rohan Rao, Sunil Patel, Stefano Pasquali, and Dhagash Mehta. HybridRAG: Integrating Knowledge Graphs and Vector Retrieval Augmented Generation for Efficient Information Extraction, August 2024.
[28] Lang Cao, Qingyu Chen, and Yue Guo. EHR-RAG: Bridging Long-Horizon Structured Electronic Health Records and Large Language Models via Enhanced Retrieval-Augmented Generation, January 2026.
[29] Hansi Zhang, Wilma Schmidt, Xiaozhi Shen, Qiushi Cao, Sebastian Monka, and Adrian Paschke. Knowledge Graph Construction towards a Graph RAG-Enhanced Intelligent Maintenance Chatbot. September 2025.
[30] Mohammad Sadeq Abolhasani and Rong Pan. Leveraging LLM for Automated Ontology Extraction and Knowledge Graph Generation, November 2024.
[31] Mojtaba Nayyeri, Athish A. Yogi, Nadeen Fathallah, Ratan Bahadur Thapa, Hans-Michael Tautenhahn, Anton Schnurpel, and Steffen Staab. Retrieval-Augmented Generation of Ontologies from Relational Databases, June 2025. arXiv:2506.01232 [cs].
[32] Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the Middle: How Language Models Use Long Contexts, November 2023. arXiv:2307.03172 [cs.CL].
[33] Weidong Bao, Yilin Wang, Ruyu Gao, Fangling Leng, Yubin Bao, and Ge Yu. DIAL-KG: Schema-Free Incremental Knowledge Graph Construction via Dynamic Schema Induction and Evolution-Intent Assessment, March 2026. arXiv:2603.20059 [cs.AI].
[34] Hongbin Ye, Honghao Gui, Xin Xu, Xi Chen, Huajun Chen, and Ningyu Zhang. Schema-adaptable Knowledge Graph Construction, November 2023. arXiv:2305.08703 [cs.CL].
[35] Bowen Zhang and Harold Soh. Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction, April 2024. arXiv:2404.03868 [cs] version: 1.
[36] Vassilis Christophides, Vasilis Efthymiou, Themis Palpanas, George Papadakis, and Kostas Stefanidis. End-to-End Entity Resolution for Big Data: A Survey, August 2020. arXiv:1905.06397 [cs.DB].
[37] Yilun Zheng, Dan Yang, Jie Li, Lin Shang, Lihui Chen, Jiahao Xu, and Sitao Luan. Less is More: Denoising Knowledge Graphs For Retrieval Augmented Generation, October 2025.
[38] Alexandros Zeakis, George Papadakis, Dimitrios Skoutas, and Manolis Koubarakis. Pre-trained Embeddings for Entity Resolution: An Experimental Analysis [Experiment, Analysis & Benchmark], April 2023. arXiv:2304.12329 [cs.DB].
[39] Daniel Obraczka, Jonathan Schuchart, and Erhard Rahm. EAGER: Embedding-Assisted Entity Resolution for Knowledge Graphs, January 2021. arXiv:2101.06126 [cs.LG].
[40] Areej Jaber and Paloma Martínez. Disambiguating Clinical Abbreviations Using a One-Fits-All Classifier Based on Deep Learning Techniques. Methods of Information in Medicine, 61(Suppl 1):e28--e34, February 2022.
[41] Sheng-Feng Sung, Ya-Han Hu, and Chong-Yan Chen. Disambiguating Clinical Abbreviations by One-to-All Classification: Algorithm Development and Validation Study. JMIR Medical Informatics, 12(1):e56955, October 2024.
[42] Kimihiro Hasegawa, Wiradee Imrattanatrai, Zhi-Qi Cheng, Masaki Asada, Susan Holm, Yuran Wang, Ken Fukuda, and Teruko Mitamura. ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding, November 2025. arXiv:2410.22211 [cs.CL] version: 2.
[43] Hexuan Deng, Wenxiang Jiao, Xuebo Liu, Min Zhang, and Zhaopeng Tu. NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates, October 2024. arXiv:2410.20814 [cs.CL].
[44] Yixuan Tang and Yi Yang. MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries, January 2024. arXiv:2401.15391 [cs.CL].
[45] Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Nikolaos Aletras and Orphee De Clercq, editors, Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, pages 150--158, St. Julians, Malta, March 2024. Association for Computational Linguistics.
[46] Nitika Mathur, Timothy Baldwin, and Trevor Cohn. Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics, June 2020.
[47] Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, August 2019.
[48] Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. BERTScore: Evaluating Text Generation with BERT, April 2019.
[49] Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, May 2023. arXiv:2303.16634 [cs].
[50] Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, May 2023.
[51] Chris Samarinas, Alexander Krubner, Alireza Salemi, Youngwoo Kim, and Hamed Zamani. Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation, May 2025. arXiv:2501.03545 [cs.CL].
[52] Yuan Sui, Mengyu Zhou, Mingjie Zhou, Shi Han, and Dongmei Zhang. Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical Study, July 2024. arXiv:2305.13062 [cs.CL].