簡易檢索 / 詳目顯示

研究生: 林宥宏
Lin, Yu-Hung
論文名稱: 基於經驗法則一致性和道德預測的道德回應生成
Moral Response Generation based on Rule of Thumb Agreement and Morality Prediction
指導教授: 吳宗憲
Wu, Chung-Hsien
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2024
畢業學年度: 112
語文別: 英文
論文頁數: 75
中文關鍵詞: 對話系統 、道德回應 、安全回應 、經驗法則
外文關鍵詞: Dialogue System, Moral Response, Safe Response, Rule of Thumb
相關次數: 點閱:124  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來,人機互動對話系統在各個領域上皆有廣泛的應用,然而,隨著人機互動對話系統變得越來越普遍,必須解決其衍生出的道德問題。在許多情況下,對話系統無法生成安全的回應,甚至可能產生或同意不安全的內容,例如不道德、粗魯或危險的訊息。
    本論文介紹了生成符合人類道德規範的回應的方法,我們除了上下文訊息外,也額外生成了相對應的經驗法則,將兩者一併交給模型,讓經驗法則達到引導的效果,我們進一步設計了一個損失函數,來衡量回應句子與經驗法則的一致性,從而使訓練後生成的回應句能夠與經驗法則保持一致。此外,我們也通過道德預測的方法,使生成的完整句子更具有道德屬性。
    我們在ProsocialDialog資料集上驗證了我們的方法,共使用了120236筆訓練資料跟25029筆測試資料,在評測上,我們使用了BLEU和PPL 兩個傳統的對話系統評測指標,以及Moral Answer Score和Agreement Score兩個道德對話系統的評測指標,結果顯示,與基準系統相比,我們的系統在道德指標上取得了顯著的改善,Moral Answer Score 提升了0.3706,Agreement Score提升了0.2718,人工評測的結果也證明,我們的系統不僅在道德性方面表現更佳,還在上下文連貫性、尊重使用者和關懷使用者等層面上達到了更好的表現。整體而言,這些結果顯示了我們方法的有效性。

    In recent years, human-computer interaction dialogue systems have seen widespread applications across various fields. However, as these systems become more prevalent, it is crucial to address the ethical issues they bring about. In many instances, dialogue systems fail to generate safe responses and may even produce or agree with unsafe content, such as unethical, rude, or dangerous messages.
    This thesis introduces methods for generating responses that comply with human moral standards. In addition to contextual information, we also generate corresponding rules of thumb and present both to the model, allowing the rules of thumb to guide the responses. We further designed a loss function to measure the agreement between the response sentences and the rules of thumb, ensuring that the responses generated after training align with the rules of thumb. Moreover, we used morality prediction methods to make the generated sentences more ethically sound.
    We validated our method on the ProsocialDialog dataset, using a total of 120,236 training instances and 25,029 test instances. For evaluation, we employed two traditional dialogue system metrics, BLEU and PPL, as well as two moral dialogue system metrics, Moral Answer Score and Agreement Score. The results showed that compared to the baseline system, our system achieved significant improvements in the moral metrics, with a 0.3706 increase in the Moral Answer Score and a 0.2718 increase in the Agreement Score. Human evaluations also demonstrated that our system not only performed better in terms of morality but also achieved better performance in contextual coherence, user respect, and user care. Overall, these results demonstrate the effectiveness of our approach.

    摘要 I Abstract III 致謝 V Contents VI List of Tables IX List of Figures XI Chapter 1 Introduction 1 1.1 Background 1 1.2 Motivation 2 1.3 Literature Review 3 1.3.1 Natural Language Generation 3 1.3.2 Task-Oriented Dialogue Systems 6 1.3.3 Chit-Chat Dialogue Systems 6 1.3.4 Moral Dialogue Systems 8 1.4 Problems 10 1.5 Brief Description of Research Methods 11 Chapter 2 Moral Dialogue Datasets 12 2.1 ProsocialDialog 12 2.2 The Moral Integrity Corpus 18 Chapter 3 Proposed Methods 24 3.1 RoT Model Training 24 3.1.1 T5 24 3.1.2 RoT Generation Model 26 3.1.3 RoBERTa 27 3.1.4 RoT Agreement Model 28 3.2 Dialogue Response Model Training 29 3.2.1 BlenderBot 29 3.2.2 Dialogue Response Model 31 3.3 Morality Prediction 32 3.3.1 LSTM 33 3.3.2 Controlled Text Generation 34 3.3.3 Morality Prediction Model 35 Chapter 4 Experimental Setup and Results 37 4.1 Evaluation Metrics 37 4.1.1 BLEU score 37 4.1.2 Rouge 39 4.1.3 Perplexity 40 4.1.4 Moral Metrics 41 4.1.5 Human Subjective Evaluation 43 4.2 Experimental Setup and Results 45 4.2.1 RoT Agreement Model 45 4.2.2 RoT Generation Model 46 4.2.3 Morality Prediction Model 48 4.2.4 Baseline Systems 49 4.2.5 Our Integrated System 50 4.2.6 Dialogue Examples 54 Chapter 5 Conclusion and Future Work 58 Reference 59

    [1] L. Laranjo, A. G. Dunn, H. L. Tong, A. B. Kocaballi, J. Chen, R. Bashir, D. Surian, B. Gallego, F. Magrabi, and A. Y. Lau, "Conversational agents in healthcare: a systematic review," Journal of the American Medical Informatics Association, vol. 25, no. 9, pp. 1248-1258, 2018.
    [2] A. N. Vaidyam, H. Wisniewski, J. D. Halamka, M. S. Kashavan, and J. B. Torous, "Chatbots and conversational agents in mental health: a review of the psychiatric landscape," The Canadian Journal of Psychiatry, vol. 64, no. 7, pp. 456-464, 2019.
    [3] R. Bavaresco, D. Silveira, E. Reis, J. Barbosa, R. Righi, C. Costa, R. Antunes, M. Gomes, C. Gatti, and M. Vanzin, "Conversational agents in business: A systematic literature review and future research directions," Computer Science Review, vol. 36, p. 100239, 2020.
    [4] G. Molnár and Z. Szüts, "The role of chatbots in formal education," in 2018 IEEE 16th International Symposium on Intelligent Systems and Informatics (SISY), 2018: IEEE, pp. 000197-000202.
    [5] S. Yang and C. Evans, "Opportunities and challenges in using AI chatbots in higher education," in Proceedings of the 2019 3rd International Conference on Education and E-Learning, 2019, pp. 79-83.
    [6] S. Laumer, C. Maier, and F. T. Gubler, "Chatbot acceptance in healthcare: Explaining user adoption of conversational agents for disease diagnosis," 2019.
    [7] J. Deng, H. Sun, Z. Zhang, J. Cheng, and M. Huang, "Recent advances towards safe, responsible, and moral dialogue systems: A survey," arXiv preprint arXiv:2302.09270, vol. 1, 2023.
    [8] D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, and K. Ndousse, "Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned," arXiv preprint arXiv:2209.07858, 2022.
    [9] P. Hu and Y. Lu, "Dual humanness and trust in conversational AI: A person-centered approach," Computers in Human Behavior, vol. 119, p. 106727, 2021.
    [10] Q. V. Liao, M. Mas-ud Hussain, P. Chandar, M. Davis, Y. Khazaeni, M. P. Crasso, D. Wang, M. Muller, N. S. Shami, and W. Geyer, "All work and no play?," in Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, 2018, pp. 1-13.
    [11] W. Wang and I. Benbasat, "Attributions of trust in decision support technologies: A study of recommendation agents for e-commerce," Journal of Management Information Systems, vol. 24, no. 4, pp. 249-273, 2008.
    [12] J. Weizenbaum, "ELIZA—a computer program for the study of natural language communication between man and machine," Communications of the ACM, vol. 9, no. 1, pp. 36-45, 1966.
    [13] Z. Liang, H. Hu, C. Xu, J. Miao, Y. He, Y. Chen, X. Geng, F. Liang, and D. Jiang, "Learning neural templates for recommender dialogue system," arXiv preprint arXiv:2109.12302, 2021.
    [14] D. E. Rumelhart, G. E. Hinton, and R. J. Williams, "Learning representations by back-propagating errors," nature, vol. 323, no. 6088, pp. 533-536, 1986.
    [15] S. Hochreiter and J. Schmidhuber, "Long short-term memory," Neural computation, vol. 9, no. 8, pp. 1735-1780, 1997.
    [16] D. Bahdanau, K. Cho, and Y. Bengio, "Neural machine translation by jointly learning to align and translate," arXiv preprint arXiv:1409.0473, 2014.
    [17] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, "Attention is all you need," Advances in neural information processing systems, vol. 30, 2017.
    [18] X. Li, Y.-N. Chen, L. Li, J. Gao, and A. Celikyilmaz, "End-to-end task-completion neural dialogue systems," arXiv preprint arXiv:1703.01008, 2017.
    [19] T. Zhao and M. Eskenazi, "Towards end-to-end learning for dialog state tracking and management using deep reinforcement learning," arXiv preprint arXiv:1606.02560, 2016.
    [20] Q. Chen, Z. Zhuo, and W. Wang, "Bert for joint intent classification and slot filling," arXiv preprint arXiv:1902.10909, 2019.
    [21] M. Forbes, J. D. Hwang, V. Shwartz, M. Sap, and Y. Choi, "Social chemistry 101: Learning to reason about social and moral norms," arXiv preprint arXiv:2011.00620, 2020.
    [22] S. Roller, E. Dinan, N. Goyal, D. Ju, M. Williamson, Y. Liu, J. Xu, M. Ott, K. Shuster, and E. M. Smith, "Recipes for building an open-domain chatbot," arXiv preprint arXiv:2004.13637, 2020.
    [23] E. M. Smith, M. Williamson, K. Shuster, J. Weston, and Y.-L. Boureau, "Can you put it all together: Evaluating conversational agents' ability to blend skills," arXiv preprint arXiv:2004.08449, 2020.
    [24] E. Dinan, S. Roller, K. Shuster, A. Fan, M. Auli, and J. Weston, "Wizard of wikipedia: Knowledge-powered conversational agents," arXiv preprint arXiv:1811.01241, 2018.
    [25] S. Zhang, E. Dinan, J. Urbanek, A. Szlam, D. Kiela, and J. Weston, "Personalizing dialogue agents: I have a dog, do you have pets too?," arXiv preprint arXiv:1801.07243, 2018.
    [26] E. Dinan, S. Humeau, B. Chintagunta, and J. Weston, "Build it break it fix it for dialogue safety: Robustness from adversarial human attack," arXiv preprint arXiv:1908.06083, 2019.
    [27] J. Xu, D. Ju, M. Li, Y.-L. Boureau, J. Weston, and E. Dinan, "Bot-adversarial dialogue for safe conversational agents," in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021, pp. 2950-2968.
    [28] H. Kim, Y. Yu, L. Jiang, X. Lu, D. Khashabi, G. Kim, Y. Choi, and M. Sap, "Prosocialdialog: A prosocial backbone for conversational agents," arXiv preprint arXiv:2205.12688, 2022.
    [29] C. Ziems, J. A. Yu, Y.-C. Wang, A. Halevy, and D. Yang, "The moral integrity corpus: A benchmark for ethical dialogue systems," arXiv preprint arXiv:2204.03021, 2022.
    [30] N. Meade, S. Gella, D. Hazarika, P. Gupta, D. Jin, S. Reddy, Y. Liu, and D. Hakkani-Tür, "Using in-context learning to improve dialogue safety," arXiv preprint arXiv:2302.00871, 2023.
    [31] S. Kim, S. Dai, M. Kachuee, S. Ray, T. Taghavi, and S. Yoon, "GrounDial: Human-norm Grounded Safe Dialog Response Generation," arXiv preprint arXiv:2402.08968, 2024.
    [32] D. Hendrycks, C. Burns, S. Basart, A. Critch, J. Li, D. Song, and J. Steinhardt, "Aligning ai with shared human values," arXiv preprint arXiv:2008.02275, 2020.
    [33] M. Sap, S. Gabriel, L. Qin, D. Jurafsky, N. A. Smith, and Y. Choi, "Social bias frames: Reasoning about social and power implications of language," arXiv preprint arXiv:1911.03891, 2019.
    [34] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, and A. Askell, "Language models are few-shot learners," Advances in neural information processing systems, vol. 33, pp. 1877-1901, 2020.
    [35] Y. Zhang, S. Sun, M. Galley, Y.-C. Chen, C. Brockett, X. Gao, J. Gao, J. Liu, and B. Dolan, "Dialogpt: Large-scale generative pre-training for conversational response generation," arXiv preprint arXiv:1911.00536, 2019.
    [36] S. Black, L. Gao, P. Wang, C. Leahy, and S. Biderman, "Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow," If you use this software, please cite it using these metadata, vol. 58, p. 2, 2021.
    [37] C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, "Exploring the limits of transfer learning with a unified text-to-text transformer," Journal of machine learning research, vol. 21, no. 140, pp. 1-67, 2020.
    [38] A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, "GLUE: A multi-task benchmark and analysis platform for natural language understanding," arXiv preprint arXiv:1804.07461, 2018.
    [39] A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, "Superglue: A stickier benchmark for general-purpose language understanding systems," Advances in neural information processing systems, vol. 32, 2019.
    [40] P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, "Squad: 100,000+ questions for machine comprehension of text," arXiv preprint arXiv:1606.05250, 2016.
    [41] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, "Roberta: A robustly optimized bert pretraining approach," arXiv preprint arXiv:1907.11692, 2019.
    [42] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "Bert: Pre-training of deep bidirectional transformers for language understanding," arXiv preprint arXiv:1810.04805, 2018.
    [43] G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy, "Race: Large-scale reading comprehension dataset from examinations," arXiv preprint arXiv:1704.04683, 2017.
    [44] K. Yang and D. Klein, "FUDGE: Controlled text generation with future discriminators," arXiv preprint arXiv:2104.05218, 2021.
    [45] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, "Bleu: a method for automatic evaluation of machine translation," in Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 2002, pp. 311-318.
    [46] C.-Y. Lin and E. Hovy, "Automatic evaluation of summaries using n-gram co-occurrence statistics," in Proceedings of the 2003 human language technology conference of the North American chapter of the association for computational linguistics, 2003, pp. 150-157.
    [47] H. Sun, Z. Zhang, F. Mi, Y. Wang, W. Liu, J. Cui, B. Wang, Q. Liu, and M. Huang, "MoralDial: A framework to train and evaluate moral dialogue systems via moral discussions," arXiv preprint arXiv:2212.10720, 2022.
    [48] T. Gao, X. Yao, and D. Chen, "Simcse: Simple contrastive learning of sentence embeddings," arXiv preprint arXiv:2104.08821, 2021.

    下載圖示
    2026-08-31公開
    QR CODE