| 研究生: |
蔡欣諭 Tsai, Hsin-Yu |
|---|---|
| 論文名稱: |
生成式人工智慧應用於台灣永續報告書之溫室氣體排放資訊擷取與組織邊界分析研究 Application of Generative Artificial Intelligence to Greenhouse Gas Emission Information Extraction and Organizational Boundary Analysis from Taiwan's ESG Reports |
| 指導教授: |
李昇暾
Li, Sheng-Tun |
| 學位類別: |
碩士 Master |
| 系所名稱: |
管理學院 - 工業與資訊管理學系 Department of Industrial and Information Management |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 76 |
| 中文關鍵詞: | 生成式人工智慧 、溫室氣體擷取 、提示工程 、ESG資訊揭露 、組織邊界 |
| 外文關鍵詞: | generative AI, greenhouse gas extraction, prompt engineering, ESG disclosure, organizational boundary |
| 相關次數: | 點閱:103 下載:2 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
企業永續資訊揭露需求日益增加,如何從篇幅龐大且非結構化格式的ESG永續報告書中系統化擷取溫室氣體排放數據,成為環境資訊分析領域的重要議題。本研究以台灣六大高碳排產業部分上市櫃公司2023年至2024年之52份永續報告書作為研究對象,針對溫室氣體範疇一、範疇二的絕對排放量及碳排放強度三項指標進行擷取與分類,建構以生成式人工智慧為核心的自動化溫室氣體資訊擷取流程。
前處理階段以關鍵詞式候選頁檢索Top-15設定縮小資料輸入範圍,在相同候選頁、任務定義與輸出欄位條件下,以規則式方法作為基準,比較生成式人工智慧零樣本提示與少樣本提示之擷取效能,同時評估碳排強度分類準確度與組織邊界準確度。
實驗結果顯示,少樣本提示之整體擷取效能及組織邊界準確度為三種方法中最佳,大型語言模型經由任務分解的提示設計與代表性範例,可將高異質性永續報告書轉為結構化溫室氣體資料,然而階層式表格判讀、跨頁欄位對應、同單位表頭及揭露邊界語意判定,仍具擷取限制。
本研究建立之候選頁檢索、欄位定義與評估框架具通用性,可作為後續泛化性處理永續報告書或其餘溫室氣體指標擷取任務之方法論基礎,提升ESG資訊擷取效率並提供實證依據。
The present study develops a generative AI-based pipeline for extracting greenhouse gas (GHG) information from Taiwan’s ESG reports. The dataset comprises 52 sustainability reports covering 2023 and 2024 from 26 companies in six carbon-intensive industries. Three extraction tasks are examined:Scope 1 absolute emissions, Scope 2 absolute emissions, and carbon intensity. A keyword-based retrieval method, supplemented by numerical density, first selects the Top-15 candidate pages from each report. A rule-based baseline, zero-shot prompting, and few-shot prompting are then applied to the same candidate pages under identical task definitions and output requirements. The ground truth dataset is established through independent annotation by two annotators, inter-annotator agreement analysis, and third-party adjudication. An extraction is considered correct only when its normalized value and unit match the ground truth. Few-shot prompting achieves the highest overall F1-score of 0.689, compared with 0.592 for zero-shot prompting and 0.255 for the rule-based baseline. Its F1-scores for Scope 1, Scope 2, and carbon intensity are 0.726, 0.693, and 0.646, respectively. Zero-shot prompting achieves the highest carbon-intensity classification accuracy, whereas few-shot prompting provides the highest organizational-boundary accuracy. The remaining errors mainly involve hierarchical tables, cross-page associations, shared unit headers, and organizational-boundary interpretation. Retaining source pages and source text supports manual verification of the extracted results. These findings show that candidate-page retrieval, combined with few-shot prompting, provides a feasible approach for automated extraction of GHG information from heterogeneous sustainability reports without model fine-tuning.
行政院環境部.(2024.2). 事業應盤查登錄及查驗溫室氣體排放量之排放源. Retrieved from https://ghgregistry.moenv.gov.tw
行政院金融監督管理委員會.(2022.3). 上市櫃公司永續發展路徑圖. Retrieved from https://www.sfb.gov.tw
Antonini, C., & Larrinaga, C. (2017). Planetary Boundaries and Sustainability Indicators. A Survey of Corporate Reporting Boundaries. Sustainable Development, 25(2), 123–137. https://doi.org/10.1002/sd.1667
Archel, P., Fernández, M., & Larrinaga, C. (2008). The Organizational and Operational Boundaries of Triple Bottom Line Reporting: A Survey. Environmental Management, 41(1), 106–117. https://doi.org/10.1007/s00267-007-9029-7
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural Machine Translation by Jointly Learning to Align and Translate. arXiv. https://doi.org/10.48550/arXiv.1409.0473
Beck, J., Steinberg, A., Dimmelmeier, A., Domenech Burin, L., Kormanyos, E., Fehr, M., & Schierholz, M. (2025). Addressing data gaps in sustainability reporting: A benchmark dataset for greenhouse gas emission extraction. Scientific Data, 12(1), 1497. https://doi.org/10.1038/s41597-025-05664-8
Bolton, P., & Kacperczyk, M. (2021). Do investors care about carbon risk? Journal of Financial Economics, 142(2), 517–549. https://doi.org/10.1016/j.jfineco.2021.05.008
Bronzini, M., Nicolini, C., Lepri, B., Passerini, A., & Staiano, J. (2024). Glitter or gold? Deriving structured insights from sustainability reports via large language models. EPJ Data Science, 13(1), 1–41. https://doi.org/10.1140/epjds/s13688-024-00481-2
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., … Amodei, D. (2020). Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems, 33, 1877–1901. https://proceedings.neurips.cc/paper_files/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html
Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low Kappa: I. the problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543–549. https://doi.org/10.1016/0895-4356(90)90158-L
Goel, A., Gueta, A., Gilon, O., Liu, C., Erell, S., Nguyen, L. H., Hao, X., Jaber, B., Reddy, S., Kartha, R., Steiner, J., Laish, I., & Feder, A. (2023). LLMs Accelerate Annotation for Medical Information Extraction. Proceedings of the 3rd Machine Learning for Health Symposium, Proceedings of Machine Learning Research, 225, 82–100. https://proceedings.mlr.press/v225/goel23a.html
Gwet, K. L. (2008). Computing inter‐rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1), 29–48. https://doi.org/10.1348/000711006X126600
Hu, Y., Chen, Q., Du, J., Peng, X., Keloth, V. K., Zuo, X., Zhou, Y., Li, Z., Jiang, X., Lu, Z., Roberts, K., & Xu, H. (2024). Improving large language models for clinical named entity recognition via prompt engineering. Journal of the American Medical Informatics Association, 31(9), 1812–1820. https://doi.org/10.1093/jamia/ocad259
Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W. (2020). Dense Passage Retrieval for Open-Domain Question Answering. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 6769–6781. https://doi.org/10.18653/v1/2020.emnlp-main.550
Klaaßen, L., & Stoll, C. (2021). Harmonizing corporate carbon footprints. Nature Communications, 12(1), 6149. https://doi.org/10.1038/s41467-021-26349-x
Kolk, A., Levy, D., & Pinkse, J. (2008). Corporate Responses in an Emerging Climate Regime: The Institutionalization and Commensuration of Carbon Disclosure. European Accounting Review, 17(4), 719–745. https://doi.org/10.1080/09638180802489121
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. https://doi.org/10.1162/tacl_a_00638
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., & Neubig, G. (2023). Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Computing Surveys, 55(9), 1–35. https://doi.org/10.1145/3560815
Lopez, I., Swaminathan, A., Vedula, K., Narayanan, S., Nateghi Haredasht, F., Ma, S. P., Liang, A. S., Tate, S., Maddali, M., Gallo, R. J., Shah, N. H., & Chen, J. H. (2025). Clinical entity augmented retrieval for clinical information extraction. Npj Digital Medicine, 8(1), 45. https://doi.org/10.1038/s41746-024-01377-1
Luo, F., Zhang, J., Wang, Q., & Yang, C. (2025). Leveraging Prompt Engineering in Large Language Models for Accelerating Chemical Research. ACS Central Science, 11(4), 511–519. https://doi.org/10.1021/acscentsci.4c01935
Maibaum, F., Kriebel, J., & Foege, J. N. (2024). Selecting textual analysis tools to classify sustainability information in corporate reporting. Decision Support Systems, 183, 114269. https://doi.org/10.1016/j.dss.2024.114269
Manning, C. D., Raghavan, P., & Schütze, H. (2009). Introduction to Information Retrieval. Cambridge University Press.
Margot, V., Geissler, C., Franco, C. de, & Monnier, B. (2021). ESG Investments: Filtering versus Machine Learning Approaches. Applied Economics and Finance, 8(2), 1–16. https://doi.org/10.11114/aef.v8i2.5097
Martín-Domingo, L., Fernandez, J. B., Efthymiou, M., & Ali, M. I. (2025). Extracting airline emission KPIs from sustainability reports using large language models (LLMs). Transportation Research Interdisciplinary Perspectives, 33, 101599. https://doi.org/10.1016/j.trip.2025.101599
Miles, S., & Ringham, K. (2020). The boundary of sustainability reporting: Evidence from the FTSE100. Accounting, Auditing & Accountability Journal, 33(2), 357–390. https://doi.org/10.1108/AAAJ-05-2018-3478
Ntinopoulos, V., Rodriguez Cetina Biefer, H., Tudorache, I., Papadopoulos, N., Odavic, D., Risteski, P., Haeussler, A., & Dzemali, O. (2025). Large language models for data extraction from unstructured and semi-structured electronic health records: A multiple model performance evaluation. BMJ Health & Care Informatics, 32(1), e101139. https://doi.org/10.1136/bmjhci-2024-101139
Parikh, P., & Penfield, J. (2024). Automatic Question Answering From Large ESG Reports. International Journal of Data Warehousing and Mining, 20(1), 352513. https://doi.org/10.4018/IJDWM.352513
Polak, M. P., & Morgan, D. (2024). Extracting accurate materials data from research papers with conversational language models and prompt engineering. Nature Communications, 15(1), 1569. https://doi.org/10.1038/s41467-024-45914-8
Saha, R., & Maji, S. G. (2025). Does environmental, social, and governance (ESG) disclosure matter for carbon intensity? Evidence from S&P 500 firms. Journal of Environmental Management, 387, 125809. https://doi.org/10.1016/j.jenvman.2025.125809
Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S., & Chadha, A. (2025). A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications (arXiv:2402.07927). arXiv. https://doi.org/10.48550/arXiv.2402.07927
Sun, Z., Satapathy, R., Guo, D., Li, B., Liu, X., Zhang, Y., Tan, C.-A., Filho, R. S., & Goh, R. S. M. (2024). Information Extraction: Unstructured to Structured for ESG Reports. 2024 IEEE International Conference on Data Mining Workshops (ICDMW), 487–495. https://doi.org/10.1109/ICDMW65004.2024.00068
Sutskever, I., Vinyals, O., & Le, Q. V. (2014). Sequence to Sequence Learning with Neural Networks. Advances in Neural Information Processing Systems, 27. https://proceedings.neurips.cc/paper_files/paper/2014/hash/5a18e133cbf9f257297f410bb7eca942-Abstract.html
Talbot, D., & Boiral, O. (2018). GHG Reporting and Impression Management: An Assessment of Sustainability Reports from the Energy Sector. Journal of Business Ethics, 147(2), 367–383. https://doi.org/10.1007/s10551-015-2979-4
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł. ukasz, & Polosukhin, I. (2017). Attention is All you Need. Advances in Neural Information Processing Systems, 30. https://proceedings.neurips.cc/paper_files/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html
Vestrelli, R., Fronzetti Colladon, A., & Pisello, A. L. (2024). When attention to climate change matters: The impact of climate risk disclosure on firm market value. Energy Policy, 185, 113938. https://doi.org/10.1016/j.enpol.2023.113938
Vijayan, A. (2023). A Prompt Engineering Approach for Structured Data Extraction from Unstructured Text Using Conversational LLMs. 2023 6th International Conference on Algorithms Computing and Artificial Intelligence, 183–189. https://doi.org/10.1145/3639631.3639663
Wang, D., Huang, Q., Jackson, M., & Gao, J. (2024). Retrieve What You Need: A Mutual Learning Framework for Open-domain Question Answering. Transactions of the Association for Computational Linguistics, 12, 247–263. https://doi.org/10.1162/tacl_a_00646
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., & Zhou, D. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems, (Vol. 35), 24824–24837.
Weston, L., Tshitoyan, V., Dagdelen, J., Kononova, O., Trewartha, A., Persson, K. A., Ceder, G., & Jain, A. (2019). Named Entity Recognition and Normalization Applied to Large-Scale Information Extraction from the Materials Science Literature. Journal of Chemical Information and Modeling, 59(9), 3692–3702. https://doi.org/10.1021/acs.jcim.9b00470
White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer-Smith, J., & Schmidt, D. C. (2023). A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT (arXiv:2302.11382). arXiv. https://doi.org/10.48550/arXiv.2302.11382
World Resources Institute, & World Business Council for Sustainable Development. (2004). The Greenhouse Gas Protocol: A corporate accounting and reporting standard.https://ghgprotocol.org/sites/default/files/standards/ghg-protocol-revised.pdf
Wu, R., Zong, H., Wu, E., Li, J., Zhou, Y., Zhang, C., Zhang, Y., Wang, J., Tang, T., & Shen, B. (2025). Improving large language models for miRNA information extraction via prompt engineering. Computer Methods and Programs in Biomedicine, 271, 109033. https://doi.org/10.1016/j.cmpb.2025.109033
Zou, Y., Shi, M., Chen, Z., Deng, Z., Lei, Z., Zeng, Z., Yang, S., Tong, H., Xiao, L., & Zhou, W. (2025). ESGReveal: An LLM-based approach for extracting structured data from ESG reports. Journal of Cleaner Production, 489, 144572. https://doi.org/10.1016/j.jclepro.2024.144572