簡易檢索 / 詳目顯示

研究生: 林奕㚬
Lin, Yi-Chun
論文名稱: TraceMem:基於模糊痕跡理論之語篇感知長期記憶框架
TraceMem:A Discourse-Aware Long-Term Memory Framework Based on Fuzzy-Trace Theory
指導教授: 吳宗憲
Wu, Chung-Hsien
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 86
中文關鍵詞: 長期對話記憶記憶增強代理對話語篇解析可追溯記憶模糊痕跡理論
外文關鍵詞: long-term conversational memory, memory-augmented agents, Fuzzy-Trace Theory, Dialogue Discourse Parsing, traceable memory
相關次數: 點閱:21下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究提出 TraceMem,一套受模糊痕跡理論啟發的語篇感知長期對話記憶框架,旨在提升大型語言模型代理對歷史互動資訊的組織、檢索與回溯能力。在跨越多輪與多個會話的長期互動中,持續累積的對話歷史容易超出模型的上下文視窗。現有記憶增強型代理多將歷史資訊表示為對話片段、摘要、事實或結構化記憶,並依據查詢相關性獨立進行檢索。然而,這類方法未必持續保留抽象記憶、來源語句與支持脈絡之間的明確連結,使檢索所得證據可能個別相關,整體卻缺乏原始對話結構。
    為解決上述問題,TraceMem 將長期對話資訊組織為互補的要旨記憶與壓縮逐字痕跡。系統首先依據語意變化將對話切分為主題一致的片段,並從各片段產生要旨記憶及識別直接支持其內容的來源語句。接著,TraceMem 透過對話語篇解析,找出與來源語句具有問答、解釋、延伸及修正等語篇關係的支持語句,建立具語篇結構的證據痕跡。每筆要旨記憶皆透過明確連結對應至其來源與支持證據,使系統能在檢索要旨後回溯相關對話脈絡,並於記憶整併過程中共同維護要旨及其證據。
    本研究於長期對話記憶基準資料集上進行評估,採用大型語言模型評審準確率、詞元層級精確率與召回率之調和平均,以及一元詞組重疊率衡量回答品質。實驗結果顯示,TraceMem 在不同生成模型與評估模型設定下皆展現良好的長期對話問答能力,整體評審準確率優於所比較的代表性記憶方法,並在自動評估指標上取得具競爭力或較佳的表現。消融實驗進一步顯示,來源連結證據、語篇關係及記憶更新均有助於提升整體效能,而證據壓縮則能在大致維持回答品質的同時降低記憶與檢索成本。研究結果說明,將抽象要旨與具語篇結構的對話證據建立持久且可回溯的連結,有助於支援長期對話代理的記憶檢索與推理。

    This study proposes TraceMem, a discourse-aware long-term conversational memory framework inspired by Fuzzy-Trace Theory. TraceMem aims to improve the ability of large language model agents to organize, retrieve, and trace information from historical interactions. In long-term interactions spanning multiple turns and sessions, continuously accumulated dialogue histories may exceed the limited context windows of language models. Existing memory-augmented agents commonly represent historical information as dialogue segments, summaries, facts, or structured memories and retrieve them independently according to their relevance to a query. However, these approaches do not necessarily maintain explicit links among abstract memories, source utterances, and supporting conversational contexts. Consequently, the retrieved evidence may be individually relevant yet collectively lack the structure of the original conversation.
    To address this problem, TraceMem organizes long-term conversational information into complementary gist memories and compressed verbatim traces. The framework first segments conversations into semantically coherent fragments, generates gist memories from these fragments, and identifies the source utterances that directly support each gist. Dialogue Discourse Parsing is then employed to identify supporting utterances connected to the source utterances through discourse relations, such as question–answer, explanation, elaboration, and correction. These utterances are organized into discourse-structured evidence traces. Each gist memory is explicitly linked to its source and supporting evidence, allowing the system to recover relevant conversational contexts after retrieving a gist and to jointly maintain the gist and its evidence during memory consolidation.
    TraceMem is evaluated on the LoCoMo long-term conversational memory benchmark using LLM-as-a-Judge accuracy, token-level F1, and BLEU-1. Experimental results demonstrate that TraceMem achieves strong long-term conversational question-answering performance across different generation and evaluation model settings. It outperforms the representative memory methods under comparison in overall LLM-as-a-Judge accuracy and achieves competitive or superior results on the automatic evaluation metrics. Ablation studies further show that source-linked evidence, discourse relations, and memory updating contribute to overall performance, while evidence compression reduces memory and retrieval costs while largely preserving answer quality. These results suggest that establishing persistent and traceable links between abstract gist memories and discourse-structured conversational evidence can effectively support memory retrieval and reasoning in long-term conversational agents.

    摘要 I Abstract III 致謝 V Content VII List of Tables X List of Figures XI Chapter 1 Introduction 1 1.1 Background 1 1.2 Motivation 3 1.3 Literature Review 5 1.3.1 Long-Term Memory for LLM-Based Agents 5 1.3.2 Memory Representation and Management 6 1.3.3 Fuzzy-Trace Theory 8 1.3.4 Dialogue Discourse Parsing 11 1.4 Problems 13 1.5 Overview of TraceMem 15 Chapter 2 Proposed Method 16 2.1 Overview of TraceMem 17 2.2 Fragment Segmentation 18 2.2.1 Dialogue Representation and Block-Level Similarity 19 2.2.2 Depth-Based Boundary Detection 21 2.2.3 Turn-Level Boundary Refinement 22 2.3 Gist Memory Generation 24 2.4 Discourse-Aware Verbatim Memory Generation 25 2.4.1 Anchor Utterance Identification 26 2.4.2 Discourse-Aware Evidence Extraction 27 2.4.3 Verbatim Memory Compression and Construction 28 2.5 Memory Retrieval and Response Generation 30 2.6 Offline Memory Consolidation 31 2.6.1 Candidate Memory Selection 32 2.6.2 Consolidation Policy Decision 33 2.6.3 Consolidation Action Execution 34 Chapter 3 Datasets 35 3.1 LoCoMo Dataset 35 3.1.1 Dataset Overview 35 3.1.2 Question-Answering Categories 36 3.2 Molweni Dataset for Dialogue Discourse Parsing 39 Chapter 4 Experimental Setup and Results 41 4.1 Evaluation Metrics 42 4.1.1 LLM-as-a-Judge 42 4.1.2 Token-Level F1 Score 43 4.1.3 BLEU-1 44 4.2 Experimental Settings 44 4.2.1 Model and Implementation Settings 44 4.2.2 Dialogue Discourse Parser Settings 45 4.2.3 Hyperparameter Settings 47 4.2.4 Baseline Methods 47 4.3 Experiment Results and Discussion 49 4.3.1 LLM-as-a-Judge Evaluation 49 4.3.2 Automatic Evaluation Results 53 4.3.3 Ablation Study 56 4.3.4 Fragment Segmentation Parameter Analysis 60 4.3.5 Effect of Retrieval Size 60 Chapter 5 Conclusion and Future Work 61 References 63 Appendix 66 A. Source-Linked Gist Memory Generation Prompt 66 B. Memory Update and Consolidation Prompt 68 C. Discourse-Aware Answer Generation Prompt 70

    [1] W. Zhong, L. Guo, Q. Gao, H. Ye, and Y. Wang, "Memorybank: Enhancing large language models with long-term memory," in Proceedings of the AAAI conference on artificial intelligence, 2024, vol. 38, no. 17, pp. 19724–19731.
    [2] C. Packer, V. Fang, S. G. Patil, K. Lin, S. Wooders, and J. E. Gonzalez, "MemGPT: towards LLMs as operating systems," 2023.
    [3] D. Wu, H. Wang, W. Yu, Y. Zhang, K.-W. Chang, and D. Yu, "Longmemeval: Benchmarking chat assistants on long-term interactive memory," arXiv preprint arXiv:2410.10813, 2024.
    [4] Y. Hu, Y. Wang, and J. McAuley, "Evaluating memory in llm agents via incremental multi-turn interactions," arXiv preprint arXiv:2507.05257, 2025.
    [5] P. Du, "Memory for autonomous llm agents: Mechanisms, evaluation, and emerging frontiers," arXiv preprint arXiv:2603.07670, 2026.
    [6] N. Asher, J. Hunter, M. Morey, B. Farah, and S. Afantenos, "Discourse structure and dialogue acts in multiparty dialogue: the STAC corpus," in Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), 2016, pp. 2721–2727.
    [7] K.-H. Lee, X. Chen, H. Furuta, J. Canny, and I. Fischer, "A human-inspired reading agent with gist memory of very long contexts," arXiv preprint arXiv:2402.09727, 2024.
    [8] W. Xu, Z. Liang, K. Mei, H. Gao, J. Tan, and Y. Zhang, "A-mem: Agentic memory for llm agents," Advances in Neural Information Processing Systems, vol. 38, pp. 17577–17604, 2026.
    [9] P. Chhikara, D. Khant, S. Aryan, T. Singh, and D. Yadav, "Mem0: Building production-ready ai agents with scalable long-term memory," arXiv preprint arXiv:2504.19413, 2025.
    [10] J. Kang, M. Ji, Z. Zhao, and T. Bai, "Memory os of ai agent," in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 25972–25981.
    [11] S. She, S. Huang, X. Wang, Y. Zhou, and J. Chen, "Exploring the factual consistency in dialogue comprehension of large language models," in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 6087–6100.
    [12] H. Li, C. Yang, A. Zhang, Y. Deng, X. Wang, and T.-S. Chua, "Hello again! llm-powered personalized agent for long-term dialogue," in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 5259–5276.
    [13] J. Jang, M. Boo, and H. Kim, "Conversation chronicles: Towards diverse temporal and relational dynamics in multi-session conversations," in Proceedings of the 2023 conference on empirical methods in natural language processing, 2023, pp. 13584–13606.
    [14] N. F. Liu et al., "Lost in the middle: How language models use long contexts," Transactions of the association for computational linguistics, vol. 12, pp. 157–173, 2024.
    [15] C.-P. Hsieh et al., "RULER: What's the real context size of your long-context language models?," arXiv preprint arXiv:2404.06654, 2024.
    [16] Y. Bai et al., "Longbench: A bilingual, multitask benchmark for long context understanding," in Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers), 2024, pp. 3119–3137.
    [17] A. Yehudai et al., "A Survey on Evaluation of LLM-based Agents," in Findings of the Association for Computational Linguistics: ACL 2026, 2026, pp. 26690–26714.
    [18] A. Maharana, D.-H. Lee, S. Tulyakov, M. Bansal, F. Barbieri, and Y. Fang, "Evaluating very long-term conversational memory of llm agents," in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 13851–13870.
    [19] J. Fang et al., "Lightmem: Lightweight and efficient memory-augmented generation," arXiv preprint arXiv:2510.18866, 2025.
    [20] Z. Tan et al., "In prospect and retrospect: Reflective memory management for long-term personalized dialogue agents," in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 8416–8439.
    [21] Z. Pan et al., "On memory construction and retrieval for personalized conversational agents," arXiv preprint arXiv:2502.05589, 2025.
    [22] C. J. Brainerd and V. F. Reyna, "Fuzzy-trace theory: Dual processes in memory, reasoning, and cognitive neuroscience," Advances in child development and behavior, vol. 28, pp. 41–100, 2002.
    [23] V. F. Reyna, "A new intuitionism: Meaning, memory, and development in fuzzy-trace theory," Judgment and Decision making, vol. 7, no. 3, pp. 332–359, 2012.
    [24] J. Li et al., "Molweni: A challenge multiparty dialogues-based machine reading comprehension dataset with discourse structure," in Proceedings of the 28th International Conference on Computational Linguistics, 2020, pp. 2642–2652.
    [25] Z. Shi and M. Huang, "A deep sequential model for discourse parsing on multi-party dialogues," in Proceedings of the AAAI Conference on Artificial Intelligence, 2019, vol. 33, no. 01, pp. 7007–7014.
    [26] C. Li, Y. Yin, and G. Carenini, "Dialogue discourse parsing as generation: A sequence-to-sequence LLM-based approach," in Proceedings of the 25th annual meeting of the special interest group on discourse and dialogue, 2024, pp. 1–14.
    [27] N. Reimers and I. Gurevych, "Sentence-bert: Sentence embeddings using siamese bert-networks," in Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), 2019, pp. 3982–3992.
    [28] W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou, "Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers," Advances in neural information processing systems, vol. 33, pp. 5776–5788, 2020.
    [29] M. A. Hearst, "Text tiling: Segmenting text into multi-paragraph subtopic passages," Computational linguistics, vol. 23, no. 1, pp. 33–64, 1997.
    [30] Z. Pan et al., "Llmlingua-2: Data distillation for efficient and faithful task-agnostic prompt compression," in Findings of the Association for Computational Linguistics: ACL 2024, 2024, pp. 963–981.
    [31] N. Asher and A. Lascarides, Logics of conversation. Cambridge University Press, 2003.
    [32] L. Zheng et al., "Judging llm-as-a-judge with mt-bench and chatbot arena," Advances in neural information processing systems, vol. 36, pp. 46595–46623, 2023.
    [33] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, "Bleu: a method for automatic evaluation of machine translation," in Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 2002, pp. 311–318.
    [34] W. Kwon et al., "Efficient memory management for large language model serving with pagedattention," in Proceedings of the 29th symposium on operating systems principles, 2023, pp. 611–626.
    [35] P. Lewis et al., "Retrieval-augmented generation for knowledge-intensive nlp tasks," Advances in neural information processing systems, vol. 33, pp. 9459–9474, 2020.
    [36] B. Xu et al., "Structmem: Structured memory for long-horizon behavior in llms," in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2026, pp. 122–146.

    QR CODE