簡易檢索 / 詳目顯示

研究生: 郭昱辰
Kuo, Yu-Chen
論文名稱: 基於階段感知與評估準則引導於情緒支持對話之偏好資料構建
Stage-Aware, Rubric-Guided Preference Data Construction for Emotional Support Conversation
指導教授: 吳宗憲
Wu, Chung-Hsien
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 89
中文關鍵詞: 情緒支持對話偏好資料構建直接偏好最佳化大型語言模型
外文關鍵詞: Emotional Support Conversation, Preference Data Construction, Direct Preference Optimization, Large Language Models
相關次數: 點閱:6下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 情緒支持對話 (Emotional Support Conversation, ESC) 旨在讓對話系統透過同理與適當的支持策略,緩解求助者的情緒困擾,在心理健康支持、社交互動與客服等場景中日益重要。ESConv 語料庫將此任務形式化,並依據 Hill 的助人技巧理論 (Helping Skills Theory) 將支持歷程劃分為三個階段: 探索 (Exploration)、安撫 (Comforting) 與行動 (Action)。近期研究以 Direct Preference Optimization (DPO) 提升 ESC 生成品質,其中 DecoupledESC 進一步將策略規劃與回應生成解耦,分別以兩個獨立的 DPO 目標最佳化,避免耦合式目標的最佳化衝突。
    然而,我們觀察到既有方法在構建偏好資料 (chosen/rejected pair) 時,從未將「階段正確性」納入挑選準則,因而完全沒有考慮對話當下所處的階段。我們在 ESConv 的驗證集上檢視 DecoupledESC 基線的階段預測,發現在可對應至三個階段的預測中,超過半數落在錯誤的階段,且錯誤同時涵蓋三個偏移方向,並非單一主導類型; 分層分析進一步顯示,只要預測階段錯誤,無論錯誤偏向哪個方向,回應在 BLEU-1、ROUGE-L 與 F1 這些重疊型自動指標上的品質都會下降。因此,錯誤的階段預測是降低回應品質的上游因素。
    為修補此缺口,本論文在 DecoupledESC 之上,提出階段感知的偏好資料構建方法,整體訓練流程與基線完全相同,貢獻著重於「如何挑選 chosen 與 rejected」。我們在 Strategy Planner (SP) 端以三方向對稱的階段偏移過採樣 (symmetric stage-drift oversampling) 強化對錯誤階段的懲罰; 在 Response Generator (RG) 端提出評估準則引導的挖掘 (rubric-guided mining),將原本用來檢核回應是否符合當前階段的評估準則,轉化為構建偏好資料的訊號,以此挑選出形成階段對比的 chosen 與 rejected 回應。
    我們在 Llama-3.1-8B-Instruct 與 Qwen2.5-7B-Instruct 兩個骨幹模型上,與 DecoupledESC 基線在相同的訓練與推論設定下比較。實驗結果顯示,在兩個骨幹模型上,本方法在 Fluency、Professionalism、Empathy 與 Helpfulness 四個 LLM 評審指標,以及 BLEU、ROUGE-L 與 F1 等重疊型自動指標上全數優於基線; 回應品質指標中僅有衡量多樣性的 Distinct-1 略為下降,這是回應長度增加帶來的副作用。同時,本方法大幅降低「過早偏移至行動階段」的比率 (ISAR),兩個骨幹模型分別下降約 8.6 與 8.2 個百分點。消融實驗確認方法中的每一個元件皆為必要,且各項超參數皆位於局部最佳點。
    本論文顯示,只要將原本用於評估的階段檢核準則轉化為構建偏好資料的訊號,便能在不更動訓練流程的前提下,提升情緒支持回應的品質。

    Emotional Support Conversation (ESC) aims to relieve a help-seeker's emotional distress through empathy and appropriate support strategies, and is increasingly important in mental-health support, social interaction, and customer service. The ESConv corpus formalizes this task and, following Hill's Helping Skills Theory, organizes the supporting process into three stages: Exploration, Comforting, and Action. Recent work improves ESC generation with Direct Preference Optimization (DPO); in particular, DecoupledESC decouples strategy planning from response generation and optimizes each with a separate DPO objective, avoiding the optimization conflict of a coupled objective.
    We observe, however, that existing preference-data construction never uses stage correctness as a criterion for selecting chosen and rejected pairs, and is therefore agnostic to the counseling-stage structure. Auditing the DecoupledESC baseline on the ESConv held-out validation set, we find that more than half of the predictions that map to one of the three stages fall into the wrong Hill stage, with errors spanning all three shift directions rather than a single dominant type. A stratified analysis further shows that a wrong predicted stage, in any direction, lowers response quality on the overlap-based automatic metrics. A wrong stage is thus an upstream factor that degrades response quality.
    To close this gap, we propose stage-aware preference-data construction as an extension of DecoupledESC. The training pipeline is unchanged from the baseline; our contribution lies in how chosen and rejected examples are selected. On the strategy-planning (SP) model we apply symmetric, three-direction stage-drift oversampling to strengthen the penalty for predicting a wrong-stage strategy. On the response-generation (RG) model we introduce rubric-guided mining, which repurposes a stage-appropriateness checklist rubric, normally an evaluation tool, into a data-construction signal, and uses it to select stage-contrastive chosen and rejected responses.
    We evaluate on two backbones, Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct, against the DecoupledESC baseline under identical training and inference settings. Our method improves all four LLM-as-judge dimensions of Fluency, Professionalism, Empathy, and Helpfulness over the baseline on both backbones, and improves every overlap-based automatic metric. Among the response-quality metrics, the only decline is a small drop in the diversity metric Distinct-1, which is a side effect of the increased response length. It also reduces the rate of incorrect stage shift to Action (ISAR) by 8.6 and 8.2 percentage points. Ablation studies confirm that each component is necessary and that every hyper-parameter corresponds to a local optimum.
    Overall, this thesis shows that recasting a stage-appropriateness rubric from an evaluation metric into a preference-data-construction signal improves emotional-support response quality without altering the training pipeline.

    摘要 I Abstract III 致謝 V Table of Contents VI List of Tables IX List of Figures XI Notation XII Chapter 1 Introduction 1 1.1 Background 1 1.2 Motivation 3 1.3 Literature Review 4 1.3.1 Dialogue Systems 4 1.3.2 Emotional Support Conversation 4 1.3.3 Large Language Models and Alignment 5 1.3.4 Preference Optimization and DPO 5 1.3.5 Preference Optimization for Emotional Support and Controllable Dialogue 6 1.4 Problem Statement 7 1.5 Brief Description of Research Methods 10 1.6 Thesis Organization 10 Chapter 2 Proposed Method: Stage-Aware Preference Data Construction 11 2.1 Preliminaries 11 2.1.1 Problem Formulation of Emotional Support Conversation 11 2.1.2 Counseling Stages and Support Strategies 12 2.1.3 Neural Response Generation with Large Language Models 14 2.1.4 Preference Optimization and Direct Preference Optimization (DPO) 14 2.1.5 The DecoupledESC Baseline 16 2.2 Method Overview 17 2.3 Strategy Planning and Stage Derivation 19 2.4 Stage-Aware Preference-Pair Construction 19 2.4.1 Strategy Planner: Symmetric Stage-Drift Oversampling 19 2.4.2 Response Generator: Why Rubric-Guided Mining? 21 2.4.3 Step 0: The 12-Question Rubric and its Decomposition 22 2.4.4 Steps 1–2: Hybrid Chosen and Stage-Contrastive Rejected 26 2.5 Training Objective: Standard Decoupled DPO 31 2.6 Training and Inference Pipeline 32 Chapter 3 Datasets and Stage-Aware Preference Data Construction 34 3.1 The ESConv Corpus 34 3.2 Candidate Response Pool from the RG-SFT Model 37 3.3 The Constructed Stage-Aware Preference Datasets 38 3.3.1 Strategy-Planner Preference Dataset 38 3.3.2 Response-Generator Preference Dataset 39 3.4 Data Quality and Ethical Considerations 41 Chapter 4 Experiments 43 4.1 Evaluation Metrics 43 4.1.1 Automatic Overlap Metrics 43 4.1.2 Diversity Metric 44 4.1.3 LLM-Based Quality Metrics 44 4.1.4 Stage-Level Metrics 45 4.2 Experimental Setup 46 4.3 Results and Discussion 48 4.3.1 Main Results 48 4.3.2 Incorrect-Stage-Shift Analysis 50 4.3.3 Ablation: Strategy-Planner Oversampling 50 4.3.4 Ablation: Response-Generator Components 51 4.3.5 Ablation: Stage-Drift Directions 52 4.3.6 Ablation: Hyper-parameter Sweeps 52 4.3.7 Case Study 54 Chapter 5 Conclusion and Future Work 56 5.1 Conclusion 56 5.2 Limitations 57 5.3 Future Work 57 References 58 Appendix A. The 12-Question Stage-Appropriateness Rubric 62 Appendix B. LLM-as-Judge Prompts (Response-Quality Evaluation) 66 Appendix C. Full Hyper-parameter Sweep Tables 72 Appendix D. Cross-Judge Validation of the LLM-Judge Results 74

    [1] S. Liu, C. Zheng, O. Demasi, S. Sabour, Y. Li, Z. Yu, Y. Jiang, and M. Huang, "Towards emotional support dialog systems," in Proc. 59th Annu. Meeting of the Association for Computational Linguistics and the 11th Int. Joint Conf. on Natural Language Processing (ACL-IJCNLP), 2021.
    [2] C. E. Hill, Helping Skills: Facilitating Exploration, Insight, and Action, 3rd ed. Washington, DC: American Psychological Association, 2009.
    [3] R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn, "Direct preference optimization: Your language model is secretly a reward model," in Advances in Neural Information Processing Systems (NeurIPS), 2023.
    [4] C. Zhang, X. Shi, X. Zhang, Y. Zhu, Y. Yang, and Y. Luo, "DecoupledESC: Enhancing emotional support generation via strategy-response decoupled preference optimization," in Findings of the Association for Computational Linguistics: EMNLP 2025, 2025.
    [5] J. Gao, M. Galley, and L. Li, "Neural approaches to conversational AI," Foundations and Trends in Information Retrieval, vol. 13, no. 2–3, 2019.
    [6] H. Rashkin, E. M. Smith, M. Li, and Y.-L. Boureau, "Towards empathetic open-domain conversation models: A new benchmark and dataset," in Proc. 57th Annu. Meeting of the Association for Computational Linguistics (ACL), 2019.
    [7] Z. Lin, A. Madotto, J. Shin, P. Xu, and P. Fung, "MoEL: Mixture of empathetic listeners," in Proc. 2019 Conf. on Empirical Methods in Natural Language Processing and the 9th Int. Joint Conf. on Natural Language Processing (EMNLP-IJCNLP), 2019.
    [8] N. Majumder, P. Hong, S. Peng, J. Lu, D. Ghosal, A. Gelbukh, R. Mihalcea, and S. Poria, "MIME: MIMicking emotions for empathetic response generation," in Proc. 2020 Conf. on Empirical Methods in Natural Language Processing (EMNLP), 2020.
    [9] S. Sabour, C. Zheng, and M. Huang, "CEM: Commonsense-aware empathetic response generation," in Proc. AAAI Conf. on Artificial Intelligence, 2022.
    [10] Q. Tu, Y. Li, J. Cui, B. Wang, J.-R. Wen, and R. Yan, "MISC: A mixed strategy-aware model integrating COMET for emotional support conversation," in Proc. 60th Annu. Meeting of the Association for Computational Linguistics (ACL), 2022.
    [11] Y. Deng, W. Zhang, Y. Yuan, and W. Lam, "Knowledge-enhanced mixed-initiative dialogue system for emotional support conversations," in Proc. 61st Annu. Meeting of the Association for Computational Linguistics (ACL), 2023.
    [12] Y. Cheng, W. Liu, W. Li, J. Wang, R. Zhao, B. Liu, X. Liang, and Y. Zheng, "Improving multi-turn emotional support dialogue generation with lookahead strategy planning," in Proc. 2022 Conf. on Empirical Methods in Natural Language Processing (EMNLP), 2022.
    [13] W. Zhao, Y. Zhao, S. Wang, and B. Qin, "TransESC: Smoothing emotional support conversation via turn-level state transition," in Findings of the Association for Computational Linguistics: ACL 2023, 2023.
    [14] J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou, "Chain-of-thought prompting elicits reasoning in large language models," in Advances in Neural Information Processing Systems (NeurIPS), 2022.
    [15] T. Zhang, X. Zhang, J. Zhao, L. Zhou, and Q. Jin, "ESCoT: Towards interpretable emotional support dialogue systems," in Proc. 62nd Annu. Meeting of the Association for Computational Linguistics (ACL), 2024.
    [16] J. Kim, C. Mok, J. Lee, H. S. Kim, and Y. Jo, "Dialogue systems for emotional support via value reinforcement," in Proc. Association for Computational Linguistics (ACL), 2025.
    [17] C. Zheng, S. Sabour, J. Wen, Z. Zhang, and M. Huang, "AugESC: Dialogue augmentation with large language models for emotional support conversation," in Findings of the Association for Computational Linguistics: ACL 2023, 2023.
    [18] Z. Zheng, L. Liao, Y. Deng, L. Qin, and L. Nie, "Self-chats from large language models make small emotional support chatbot better," in Proc. 62nd Annu. Meeting of the Association for Computational Linguistics (ACL), 2024.
    [19] A. Grattafiori et al., "The Llama 3 herd of models," arXiv preprint arXiv:2407.21783, 2024.
    [20] Qwen Team, "Qwen2.5 technical report," arXiv preprint arXiv:2412.15115, 2024.
    [21] L. Ouyang et al., "Training language models to follow instructions with human feedback," in Advances in Neural Information Processing Systems (NeurIPS), 2022.
    [22] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, "Proximal policy optimization algorithms," arXiv preprint arXiv:1707.06347, 2017.
    [23] M. Gheshlaghi Azar, M. Rowland, B. Piot, D. Guo, D. Calandriello, M. Valko, and R. Munos, "A general theoretical paradigm to understand learning from human preferences," in Proc. 27th Int. Conf. on Artificial Intelligence and Statistics (AISTATS), 2024.
    [24] K. Ethayarajh, W. Xu, N. Muennighoff, D. Jurafsky, and D. Kiela, "Model alignment as prospect theoretic optimization," in Proc. 41st Int. Conf. on Machine Learning (ICML), 2024.
    [25] Y. Meng, M. Xia, and D. Chen, "SimPO: Simple preference optimization with a reference-free reward," in Advances in Neural Information Processing Systems (NeurIPS), 2024.
    [26] A. Kong, W. Ma, S. Zhao, Y. Li, Y. Wu, K. Wang, X. Liu, Q. Li, Y. Qin, and F. Huang, "SDPO: Segment-level direct preference optimization for social agents," in Proc. 63rd Annu. Meeting of the Association for Computational Linguistics (ACL), 2025.
    [27] R. Park, R. Rafailov, S. Ermon, and C. Finn, "Disentangling length from quality in direct preference optimization," in Findings of the Association for Computational Linguistics: ACL 2024, 2024.
    [28] W. Zhao et al., "Chain of strategy optimization makes large language models better emotional supporter," in Findings of the Association for Computational Linguistics: EMNLP 2025, 2025.
    [29] V. Viswanathan, Y. Sun, S. Ma, X. Kong, M. Cao, G. Neubig, and T. Wu, "Checklists are better than reward models for aligning language models," in Advances in Neural Information Processing Systems (NeurIPS), 2025.
    [30] V. Gallego, "Configurable preference tuning with rubric-guided synthetic data," arXiv preprint arXiv:2506.11702, 2025.
    [31] J. Chang, K.-Y. Chen, and C.-H. Wu, "Applying reinforcement learning and multi-generators for stage transition in an emotional support dialogue system," in Proc. Interspeech 2024, 2024, pp. 3545–3549.
    [32] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, "Attention is all you need," in Advances in Neural Information Processing Systems (NeurIPS), 2017.
    [33] Y. Lee, J. Kim, J. Kim, H. Cho, J. Kang, P. Kang, and N. Kim, "CheckEval: A reliable LLM-as-a-judge framework for evaluating text generation using checklists," in Proc. 2025 Conf. on Empirical Methods in Natural Language Processing (EMNLP), 2025.
    [34] N. Madani and R. K. Srihari, "ESC-Judge: A framework for comparing emotional support conversational agents," in Proc. 2025 Conf. on Empirical Methods in Natural Language Processing (EMNLP), 2025.
    [35] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, "BLEU: a method for automatic evaluation of machine translation," in Proc. 40th Annu. Meeting of the Association for Computational Linguistics (ACL), 2002.
    [36] C.-Y. Lin, "ROUGE: A package for automatic evaluation of summaries," in Text Summarization Branches Out, ACL Workshop, 2004.
    [37] J. Li, M. Galley, C. Brockett, J. Gao, and B. Dolan, "A diversity-promoting objective function for neural conversation models," in Proc. NAACL-HLT, 2016.
    [38] L. Zheng et al., "Judging LLM-as-a-judge with MT-Bench and Chatbot Arena," in Advances in Neural Information Processing Systems (NeurIPS), 2023.
    [39] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, "LoRA: Low-rank adaptation of large language models," in Int. Conf. on Learning Representations (ICLR), 2022.
    [40] Y. Zheng, R. Zhang, J. Zhang, Y. Ye, and Z. Luo, "LlamaFactory: Unified efficient fine-tuning of 100+ language models," in Proc. 62nd Annu. Meeting of the Association for Computational Linguistics (ACL), System Demonstrations, 2024.
    [41] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, "Efficient memory management for large language model serving with PagedAttention," in Proc. 29th ACM Symp. on Operating Systems Principles (SOSP), 2023.
    [42] Google, "Gemini models," Google AI for Developers, 2026. [Online]. Available: https://ai.google.dev/gemini-api/docs/models. Accessed: Aug. 2, 2026.

    QR CODE