| 研究生: |
洪啟貿 Hung, Chi-Mau |
|---|---|
| 論文名稱: |
應用檢索擴增生成於EA-MUStARD++之多模態諷刺偵測與解釋 RAG-based Multimodal Sarcasm Detection and Explanation on EA-MUStARD++ |
| 指導教授: |
吳宗憲
Wu, Chung-Hsien |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 79 |
| 中文關鍵詞: | 諷刺偵測 、諷刺解釋 、多模態 、檢索擴增生成 |
| 外文關鍵詞: | sarcasm detection, sarcasm explanation, multimodal, retrieval-augmented generation |
| 相關次數: | 點閱:17 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
近年來,多模態人工智慧(Multimodal Artificial Intelligence)與大型語言模型(Large Language Models, LLMs)的快速發展,使機器得以同時理解文字、語音與影像等不同模態資訊。然而,諷刺語句的理解除了依賴文字內容外,往往還需結合語音語調、臉部表情及視覺情境等多種線索,因此一直是多模態理解中相當具有挑戰性的研究議題。近年已有研究開始同時探討多模態諷刺偵測(Multimodal Sarcasm Detection)與諷刺解釋(Multimodal Sarcasm Explanation),希望透過自然語言解釋提升模型的可解釋性。然而,現有方法大多僅根據單一輸入樣本進行推論,缺乏可供參考的外部資訊;當樣本中的諷刺線索較弱或部分模態資訊不足時,容易產生不完整或不準確的解釋,進而影響諷刺偵測的結果。
本研究提出一套結合檢索擴增生成(Retrieval-Augmented Generation, RAG)的多模態諷刺偵測與解釋架構,以改善上述問題。首先,以 MUStARD++為基礎建立具有人工驗證解釋及模態觸發標註的 EA-MUStARD++(Explanation Annotated MUStARD++)資料集,並透過大型語言模型產生各模態描述與整體諷刺解釋,再經人工檢查與修正,以建立高品質的解釋資料。接著,以 MuVaC 為基礎架構建立多模態資訊資料庫,將每筆資料的多模態表示、人工驗證解釋及諷刺標籤儲存於資料庫中。在模型推論階段,利用輸入樣本檢索語意最相近的多模態案例,並將檢索結果及其人工驗證解釋作為外部先驗知識提供給解釋生成模型,以錨定解釋生成,避免模型在缺乏依據時產生無根據的內容;更可靠的解釋再透過潛在因果特徵支撐因果式諷刺偵測模型,提升模型在弱線索情境下的穩定性。
本研究採用分層 9:1 訓練/測試切分進行實驗,並重複執行十次,以取得平均值與標準差;使用 Accuracy、Precision、Recall 與 F1-score 評估諷刺偵測效能,並以 BLEU-1、ROUGE-1及METEOR 評估解釋品質。實驗結果顯示,所提出的方法在諷刺偵測上優於未採用檢索機制的基準模型,其中 F1-score 由80.79%提升至81.78%,Accuracy 由80.95%提升至81.95%,且各次執行間的標準差明顯降低,顯示模型穩定性亦獲改善。消融實驗結果亦證實,透過檢索相似案例及其解釋,可有效提供額外的推理依據,提升模型對多模態諷刺的理解能力與穩定性。此外,本研究使用外部 WITS 資料集進行檢索模組的相對消融實驗,其中以解釋品質的提升最為明顯。
Recent progress in multimodal artificial intelligence and large language models (LLMs) has improved the ability of machines to process text, audio, and vision together. Sarcasm nonetheless remains hard to recognize, because sarcastic intent is usually carried by subtle inconsistencies across modalities and only rarely by the wording alone. Recent studies have added natural-language explanations to multimodal sarcasm detection in order to make the models more interpretable, yet these methods still reason from the input instance and nothing else. When the sarcastic cues in that instance are weak or incomplete, the generated explanation is often unreliable, and an unreliable explanation can in turn drag detection down with it.
This thesis proposes a retrieval-augmented framework for multimodal sarcasm detection and explanation. We first reconstruct an explanation-annotated dataset, EA-MUStARD++ (Explanation Annotated MUStARD++), from MUStARD++, in which every dialogue carries modality-specific descriptions, a human-verified explanation, and modality-trigger annotations. On top of it we build a multimodal information database that stores multimodal representations together with their explanations and sarcasm labels. At inference time the framework retrieves semantically similar multimodal examples from this database and passes the retrieved representations and their human-verified explanations to the explanation generator, giving it external evidence to lean on instead of decoding from the query alone. The improved explanation then supports the causal-informed detector through the latent causal feature. The aim of the retrieved evidence is to raise explanation quality and to make sarcasm detection more robust.
Experiments use a stratified 9:1 train/test split repeated over ten runs, with results reported as mean ± standard deviation. Detection is evaluated by Accuracy, Precision, Recall, and F1-score, and explanation quality by BLEU-1, ROUGE-1, and METEOR. The proposed framework outperforms the retrieval-free baseline in detection, raising F1 from 80.79% to 81.78% and accuracy from 80.95% to 81.95%, and it roughly halves the standard deviation across runs, which indicates a more stable model. Ablation studies show that retrieving similar multimodal examples together with their explanations supplies useful external reasoning evidence, yielding more robust sarcasm representations and markedly more stable detection across runs, with the clearest explanation-quality gains observed on the external WITS dataset.
[1] S. Castro, D. Hazarika, V. Pérez-Rosas, R. Zimmermann, R. Mihalcea, and S. Poria, “Towards Multimodal Sarcasm Detection (An Obviously Perfect Paper),” in Proc. 57th Annu. Meeting Assoc. Comput. Linguistics (ACL), 2019, pp. 4619–4629.
[2] A. Ray, S. Mishra, A. Nunna, and P. Bhattacharyya, “A Multimodal Corpus for Emotion Recognition in Sarcasm,” in Proc. 13th Lang. Resources Eval. Conf. (LREC), 2022, pp. 6992–7003.
[3] S. Kumar, A. Kulkarni, M. S. Akhtar, and T. Chakraborty, “When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues,” in Proc. 60th Annu. Meeting Assoc. Comput. Linguistics (ACL), 2022, pp. 5956–5968.
[4] D. Guo et al., “MuVaC: A Variational Causal Framework for Multimodal Sarcasm Understanding in Dialogues,” in Proc. ACM Web Conf. (WWW), 2026. (arXiv:2601.20451)
[5] M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,” in Proc. 58th Annu. Meeting Assoc. Comput. Linguistics (ACL), 2020, pp. 7871–7880.
[6] A. Radford et al., “Learning Transferable Visual Models From Natural Language Supervision,” in Proc. 38th Int. Conf. Mach. Learn. (ICML), 2021, pp. 8748–8763.
[7] Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP), 2023, pp. 1–5.
[8] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in Proc. 31st Conf. Neural Inf. Process. Syst. (NeurIPS), 2017, pp. 5998–6008.
[9] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela, “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” in Proc. 34th Conf. Neural Inf. Process. Syst. (NeurIPS), 2020, pp. 9459–9474.
[10] C.-Y. Hsu, “Multimodal Sarcasm Detection Based on the Relation in Sarcasm Types on MUStARD++,” M.S. thesis, Dept. Comput. Sci. Inf. Eng., National Cheng Kung Univ., Tainan, Taiwan, 2024.
[11] A. Saha, V. Suresh, T. Hospedales, and V. Demberg, “MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection,” arXiv preprint arXiv:2510.23727, 2025.
[12] O. Khattab and M. Zaharia, “ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT,” in Proc. 43rd Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval (SIGIR), 2020, pp. 39–48.
[13] A. Reddy et al., “Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025. (arXiv:2503.19009)
[14] Y. Zhang, B. Li, H. Liu, Y. J. Lee, L. Gui, D. Fu, J. Feng, Z. Liu, and C. Li, “LLaVA-NeXT: A Strong Zero-shot Video Understanding Model,” Apr. 2024. [Online]. Available: https://llava-vl.github.io/blog/2024-04-30-llava-next-video/
[15] Y. Chu et al., “Qwen2-Audio Technical Report,” arXiv preprint arXiv:2407.10759, 2024.
[16] OpenAI, “GPT-4o mini: advancing cost-efficient intelligence,” 2024. [Online]. Available: https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/
[17] G. Jocher, A. Chaurasia, and J. Qiu, “Ultralytics YOLOv8,” 2023. [Online]. Available: https://github.com/ultralytics/ultralytics
[18] S. Robertson and H. Zaragoza, “The Probabilistic Relevance Framework: BM25 and Beyond,” Found. Trends Inf. Retrieval, vol. 3, no. 4, pp. 333–389, 2009.
[19] V. Karpukhin, B. Oğuz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih, “Dense Passage Retrieval for Open-Domain Question Answering,” in Proc. 2020 Conf. Empirical Methods Natural Lang. Process. (EMNLP), 2020, pp. 6769–6781.
[20] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,” in Proc. 2019 Conf. Empirical Methods Natural Lang. Process. 9th Int. Joint Conf. Natural Lang. Process. (EMNLP-IJCNLP), 2019, pp. 3982–3992.
[21] R. Nogueira and K. Cho, “Passage Re-ranking with BERT,” arXiv preprint arXiv:1901.04085, 2019.
[22] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: a Method for Automatic Evaluation of Machine Translation,” in Proc. 40th Annu. Meeting Assoc. Comput. Linguistics (ACL), 2002, pp. 311–318.
[23] C.-Y. Lin, “ROUGE: A Package for Automatic Evaluation of Summaries,” in Text Summarization Branches Out (ACL Workshop), 2004, pp. 74–81.
[24] S. Banerjee and A. Lavie, “METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,” in Proc. ACL Workshop Intrinsic Extrinsic Eval. Measures MT/Summarization, 2005, pp. 65–72.
[25] I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” in Proc. Int. Conf. Learn. Representations (ICLR), 2019.