簡易檢索 / 詳目顯示

研究生: 余函庭
Yu, Han-Ting
論文名稱: 由道德機器實驗評估大型語言模型的文化多元性
Evaluating Cultural Diversity of Large Language Models in the Moral Machine Experiment
指導教授: 李韶曼
Lee, Shao-Man
學位類別: 碩士
Master
系所名稱: 敏求智慧運算學院 - 智慧科技系統碩士學位學程
MS Degree Program on Intelligent Technology Systems
論文出版年: 2024
畢業學年度: 112
語文別: 英文
論文頁數: 82
中文關鍵詞: 大型語言模型文化多元性交織性道德決策道德機器實驗
外文關鍵詞: multilingual large language models, cultural diversity, intersectionality, moral decision-making, the Moral Machine Experiment
ORCID: https://orcid.org/0009-0006-8774-7247
相關次數: 點閱:163下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究針對主要大型語言模型(包括 GPT-3.5、GPT-4、Llama-2-7B、Llama-3-8B、Aquila-7B、ChatGLM-6B、ChatGLM-3-6B、Mistral-7B、Gemma-7B、BERT、RoBERTa、DistilBERT、ALBERT)進行全面的評估,透過比較不同文化區域的模型表現以及分析交叉人口統計特徵,探討不同模型在模擬人類道德決策中的文化多元性。研究中採用 k-means 集群方法對包含了 130 個國家資料的道德機器實驗資料集進行平衡且多元的抽樣,以克服之前研究方法的限制。研究結果揭示了一種階層式關係,大型語言模型在模擬較廣泛的文化群集時,相較於模擬較精細的區域或國家,表現出更高的文化多元性。換言之,儘管大型語言模型能廣泛地模擬不同文化集群的觀點,但在捕捉複雜的文化細節方面仍具有挑戰性。值得注意的是,擁有多元化預訓練語料庫的 ChatGLM 系列模型,在文化多元性方面表現特別突出,顯示出多元化的預訓練資料可以增強大型語言模型對於文化特徵的學習效果。

    然而,本研究的分析結果亦顯示出文化中的既有偏見可能會邊緣化保守的婦女及少數族群的觀點,而只反映出比較激進和以男性為主導地位的觀點。這種不平衡現象限制了觀點的表達,也損害了政策制定的公平性和全面性。此外,在分析結果中也發現有關俄羅斯-烏克蘭戰爭的描述往往存在資訊錯誤的風險。本研究通過評估大型語言模型的文化多元性並探討可能的改進方向,強調在模型開發過程中應謹慎考慮不同文化的特性,以確保能遵循倫理與責任的規範,並提升文化包容性,同時減少文化偏見和資訊錯誤的風險。

    This study conducts a comprehensive evaluation of prominent LLMs' (GPT-3.5, GPT-4, Llama-2-7B, Llama-3-8B, Aquila-7B, ChatGLM-6B, ChatGLM-3-6B, Mistral-7B, Gemma-7B, BERT, RoBERTa, DistilBERT, ALBERT) cultural diversity in simulating human moral decision-making across cultural dimensions and intersectional demographics. Employing k-means clustering for balanced, diverse country sampling from the 130-nation Moral Machine Experiment dataset addresses previous methods' limitations. Findings reveal a hierarchical relationship, with LLMs exhibiting higher cultural diversity simulating broad cultural clusters versus finer-grained zones and country-level nuances, underscoring challenges capturing intricate cultural nuances despite broader proficiency. Notably, ChatGLM-based models, with their diverse pre-training corpus, outperform in cultural diversity, suggesting diverse pre-training enhances cultural representation.

    However, the analysis unveils risks of cultural biases marginalizing conservative women and minority viewpoints across cultures, reflecting progressive and male perspective dominance. Such imbalances limit viewpoint expression, compromising policy development fairness and comprehensiveness. Furthermore, misinformation risks are identified regarding the Russian-Ukrainian war's portrayal. By evaluating LLMs' cultural diversity and identifying improvement areas, this study emphasizes carefully considering cultural representation during development to ensure ethical, responsible deployment fostering inclusivity while mitigating cultural bias and misinformation risks.

    中文摘要 i Abstract ii Acknowledgements iv Contents v List of Tables viii List of Figures ix 1 Introduction 1 2 Related Works 4 2.1 Datasets 5 2.2 Sampling Limitations 7 2.3 Evaluation 7 3 Methods 10 3.1 The Moral Machine Experiment Dataset 11 3.2 Sampling Strategy 15 3.2.1 Dataset 1 - A dataset consistent with MIT’s distribution of country samples 17 3.2.2 Dataset 2 - A dataset with equal sample sizes across countries 19 3.3 Models 19 3.3.1 Pretraining Data Sources 21 3.4 Prompts 21 3.4.1 Prompts without Persona 22 3.4.2 Prompts with Reasoning Request 23 3.4.3 Prompts in Different Languages 23 3.5 Analysis Process 24 3.5.1 Evaluation Metrics 25 3.5.2 Cultural and Demographic Dimensions 26 4 Results 30 4.1 Cultural Alignment and Diversity 31 4.1.1 Cultural Cluster 31 4.1.2 Cultural Zone 32 4.1.3 Country 33 4.1.4 Language 34 4.2 Demographic Dimensions 36 4.2.1 All Intersectional Groups 36 4.2.2 Intersectional Demographic Groups under Cultural Zones 37 5 Discussion 54 5.1 Balancing Accuracy and Cultural Diversity in LLMs 54 5.2 Cultural Granularity and LLM Performance 55 5.3 Cultural Zone Performance Variability 55 5.4 Performance Comparison of Models with Similar Configurations 56 5.5 Potential Risks and Implications 57 5.5.1 Geopolitical Tensions and Information Warfare 58 5.5.2 Political Orientation and LLM Performance 58 5.5.3 Gender Bias across Cultural Zones 58 5.6 Evolution of LLM’s Decision-Making Over Time 59 6 Conclusion 61 References 63

    [1] Abdulhai, M., Serapio-Garcia, G., Crepy, C., Valter, D., Canny, J., and Jaques, N. Moral foundations of large language models. arXiv preprint arXiv:2310.15337 (2023).
    [2] Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023).
    [3] Aher, G. V., Arriaga, R. I., and Kalai, A. T. Using large language models to simulate multiple humans and replicate human subject studies. In International Conference on Machine Learning (2023), PMLR, pp. 337–371.
    [4] Almeida, G. F., Nunes, J. L., Engelmann, N., Wiegmann, A., and de Araujo, M. ´ Exploring the psychology of llms’ moral and legal reasoning. Artificial Intelligence 333 (2024), 104145.
    [5] Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., and Wingate, D. Out of one, many: Using language models to simulate human samples. Political Analysis 31, 3 (2023), 337–351.
    [6] Arora, A., Kaffee, L.-A., and Augenstein, I. Probing pre-trained language models for cross-cultural differences in values. arXiv preprint arXiv:2203.13722 (2022).
    [7] Awad, E., Dsouza, S., Kim, R., Schulz, J., Henrich, J., Shariff, A., Bonnefon, J.-F., and Rahwan, I. The moral machine experiment. Nature 563, 7729 (2018), 59–64.
    [8] Bowleg, L. The problem with the phrase women and minorities: intersectionality—an important theoretical framework for public health. American journal of public health 102, 7 (2012), 1267–1273.
    [9] Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901.
    [10] Caliskan, A., Bryson, J. J., and Narayanan, A. Semantics derived automatically from language corpora contain human-like biases. Science 356, 6334 (2017), 183–186.
    [11] Cao, Y., Zhou, L., Lee, S., Cabello, L., Chen, M., and Hershcovich, D. Assessing cross-cultural alignment between chatgpt and human societies: An empirical study. arXiv preprint arXiv:2303.17466 (2023).
    [12] Chen, L., Zaharia, M., and Zou, J. How is chatgpt’s behavior changing over time? arXiv preprint arXiv:2307.09009 (2023).
    [13] Cheng, M., Durmus, E., and Jurafsky, D. Marked personas: Using natural language prompts to measure stereotypes in language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (2023), pp. 1504–1532.
    [14] Coates, A., and Ng, A. Y. Learning feature representations with k-means. In Neural Networks: Tricks of the Trade: Second Edition. Springer, 2012, pp. 561– 580.
    [15] Coda-Forno, J., Binz, M., Akata, Z., Botvinick, M., Wang, J., and Schulz, E. Meta-in-context learning in large language models. Advances in Neural Information Processing Systems 36 (2023), 65189–65201.
    [16] Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pretraining of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018).
    [17] Dominguez-Olmedo, R., Hardt, M., and Mendler-Dunner, C. ¨ Questioning the survey responses of large language models. arXiv preprint arXiv:2306.07951 (2023).
    [18] Du, H., Teng, S., Chen, H., Ma, J., Wang, X., Gou, C., Li, B., Ma, S., Miao, Q., Na, X., et al. Chat with chatgpt on intelligent vehicles: An ieee tiv perspective. IEEE Transactions on Intelligent Vehicles (2023).
    [19] Durmus, E., Nyugen, K., Liao, T. I., Schiefer, N., Askell, A., Bakhtin, A., Chen, C., Hatfield-Dodds, Z., Hernandez, D., Joseph, N., et al. Towards measuring the representation of subjective global opinions in language models. arXiv preprint arXiv:2306.16388 (2023).
    [20] Ferrara, E. Should chatgpt be biased? challenges and risks of bias in large language models. arXiv preprint arXiv:2304.03738 (2023).
    [21] Gao, Y., Tong, W., Wu, E. Q., Chen, W., Zhu, G., and Wang, F.-Y. Chat with chatgpt on interactive engines for intelligent driving. IEEE Transactions on Intelligent Vehicles (2023).
    [22] Geert, H. Culture’s consequences: International differences in work-related values. Beverly Hills: Sage (1980).
    [23] Gohar, U., and Cheng, L. A survey on intersectional fairness in machine learning: Notions, mitigation, and challenges. arXiv preprint arXiv:2305.06969 (2023).
    [24] Graham, J., Nosek, B. A., Haidt, J., Iyer, R., Koleva, S., and Ditto, P. H. Mapping the moral domain. Journal of personality and social psychology 101, 2 (2011), 366.
    [25] Hammerl, K., Deiseroth, B., Schramowski, P., Libovick ¨ y, J., Fraser, ` A., and Kersting, K. Do multilingual language models capture differing moral norms? arXiv preprint arXiv:2203.09904 (2022).
    [26] Hammerl, K., Deiseroth, B., Schramowski, P., Libovick ¨ y, J., ` Rothkopf, C. A., Fraser, A., and Kersting, K. Speaking multiple languages affects the moral bias of language models. arXiv preprint arXiv:2211.07733 (2022).
    [27] Henrich, J. Culture and social behavior. Current opinion in behavioral sciences 3 (2015), 84–89.
    [28] Hwang, E., Majumder, B. P., and Tandon, N. Aligning language models to user opinions. arXiv preprint arXiv:2305.14929 (2023).
    [29] Inglehart, R., and Welzel, C. Modernization, cultural change, and democracy: The human development sequence, vol. 333. Cambridge university press Cambridge, 2005.
    [30] Islam, R., Keya, K. N., Pan, S., Sarwate, A. D., and Foulds, J. R. Differential fairness: an intersectional framework for fair ai. Entropy 25, 4 (2023), 660.
    [31] Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825 (2023).
    [32] Kirk, H. R., Jun, Y., Volpin, F., Iqbal, H., Benussi, E., Dreyer, F., Shtedritski, A., and Asano, Y. Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models. Advances in neural information processing systems 34 (2021), 2611–2624.
    [33] Kramsch, C. Language and culture. AILA review 27, 1 (2014), 30–55.
    [34] Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942 (2019).
    [35] Lee, S., Peng, T.-Q., Goldberg, M. H., Rosenthal, S. A., Kotcher, J. E., Maibach, E. W., and Leiserowitz, A. Can large language models capture public opinion about global warming? an empirical assessment of algorithmic fidelity and bias. arXiv preprint arXiv:2311.00217 (2023).
    [36] Leong, W. Q., Ngui, J. G., Susanto, Y., Rengarajan, H., Sarveswaran, K., and Tjhi, W. C. Bhasa: A holistic southeast asian linguistic and cultural evaluation suite for large language models. arXiv preprint arXiv:2309.06085 (2023).
    [37] Li, C., Chen, M., Wang, J., Sitaram, S., and Xie, X. Culturellm: Incorporating cultural differences into large language models. arXiv preprint arXiv:2402.10946 (2024).
    [38] Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110 (2022).
    [39] Liang, P. P., Wu, C., Morency, L.-P., and Salakhutdinov, R. Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning (2021), PMLR, pp. 6565–6576.
    [40] Liu, Z., Lin, W., Shi, Y., and Zhao, J. A robustly optimized bert pretraining approach with post-training. In China National Conference on Chinese Computational Linguistics (2021), Springer, pp. 471–484.
    [41] Ma, W., Chiang, B., Wu, T., Wang, L., and Vosoughi, S. Intersectional stereotypes in large language models: Dataset and analysis. In Findings of the Association for Computational Linguistics: EMNLP 2023 (2023), pp. 8589–8597.
    [42] MacQueen, J., et al. Some methods for classification and analysis of multivariate observations. In Proceedings of the fifth Berkeley symposium on mathematical statistics and probability (1967), vol. 1, Oakland, CA, USA, pp. 281–297.
    [43] Milgram, S. Behavioral study of obedience. OCCUPATIONS 20 (2015), 29.
    [44] Moussa¨ıd, M., Kammer, J. E., Analytis, P. P., and Neth, H. ¨ Social influence and the collective dynamics of opinion formation. PloS one 8, 11 (2013), e78433.
    [45] Narvaez, D., Getz, I., Rest, J. R., and Thoma, S. J. Individual moral judgment and cultural ideologies. Developmental psychology 35, 2 (1999), 478.
    [46] Omrani Sabbaghi, S., Wolfe, R., and Caliskan, A. Evaluating biased attitude associations of language models in an intersectional context. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (2023), pp. 542–553.
    [47] Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback, 2022.
    [48] OZKURT, C. ¨ Comparative analysis of state-of-the-art q\&a models: Bert, roberta, distilbert, and albert on squad v2 dataset. Research Square preprint (2024).
    [49] Patson, N. D., Darowski, E. S., Moon, N., and Ferreira, F. Lingering misinterpretations in garden-path sentences: evidence from a paraphrasing task. Journal of Experimental Psychology: Learning, Memory, and Cognition 35, 1 (2009), 280.
    [50] Ramezani, A., and Xu, Y. Knowledge of cultural moral norms in large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (2023), pp. 428–446.
    [51] Ronen, S., and Shenkar, O. Mapping world cultures: Cluster formation, sources and implications. Journal of International Business Studies 44 (2013), 867–897.
    [52] Rosenbusch, H., Stevenson, C. E., and van der Maas, H. L. How accurate are gpt-3’s hypotheses about social science phenomena? Digital Society 2, 2 (2023), 26.
    [53] Rosette, A. S., de Leon, R. P., Koval, C. Z., and Harrison, D. A. Intersectionality: Connecting experiences of gender with race at work. Research in Organizational Behavior 38 (2018), 1–22.
    [54] Salewski, L., Alaniz, S., Rio-Torto, I., Schulz, E., and Akata, Z. In-context impersonation reveals large language models’ strengths and biases. Advances in Neural Information Processing Systems 36 (2024).
    [55] Sanh, V., Debut, L., Chaumond, J., and Wolf, T. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108 (2019).
    [56] Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., and Hashimoto, T. Whose opinions do language models reflect? In International Conference on Machine Learning (2023), PMLR, pp. 29971–30004.
    [57] Takemoto, K. The moral machine experiment on large language models. Royal Society Open Science 11, 2 (2024), 231393.
    [58] Tao, Y., Viberg, O., Baker, R. S., and Kizilcec, R. F. Auditing and mitigating cultural bias in llms. arXiv preprint arXiv:2311.14096 (2023).
    [59] Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Riviere, M., Kale, M. S., Love, J., et al. ` Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295 (2024).
    [60] Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023).
    [61] Wang, A., Morgenstern, J., and Dickerson, J. P. Large language models cannot replace human participants because they cannot portray identity groups. arXiv preprint arXiv:2402.01908 (2024).
    [62] Wang, W., Jiao, W., Huang, J., Dai, R., Huang, J.-t., Tu, Z., and Lyu, M. R. Not all countries celebrate thanksgiving: On the cultural dominance in large language models. arXiv preprint arXiv:2310.12481 (2023).
    [63] Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., Xia, X., et al. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414 (2022).
    [64] Ziems, C., Held, W., Shaikh, O., Chen, J., Zhang, Z., and Yang, D. Can large language models transform computational social science? arxiv. Preprint posted online on April 12 (2023).

    下載圖示
    2026-08-16公開
    QR CODE