簡易檢索 / 詳目顯示

研究生: 許嘉芸
Hsu, Chia-Yun
論文名稱: 基於MUStARD++資料庫中諷刺類型關係之多模態諷刺偵測
Multimodal Sarcasm Detection Based on the Relation in Sarcasm Types on MUStARD++
指導教授: 吳宗憲
Wu, Chung-Hsien
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2024
畢業學年度: 112
語文別: 英文
論文頁數: 72
中文關鍵詞: 諷刺偵測 、諷刺類型 、多模態 、多任務學習
外文關鍵詞: sarcasm detection, sarcasm types, multimodal, multitask learning
相關次數: 點閱:109  下載:1 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來,人們對於諷刺偵測系統的興趣日益濃厚,這源於人們渴望識別與理解人類表達的複雜情感,在對話機器人、電子商務等自動回話系統越來越普及的時空背景下,如何正確理解人們表達的意思至關重要,錯誤的理解可能會導致錯誤的回話。
    目前研究大多做在單模態上,且沒有考慮到語言學中諷刺不同類型的特性,語言學將諷刺分為四種類型,不同類型的說話者意圖、特徵都不相同,需要考慮不同的資訊,且需要考慮文字以外的模態,因此我們針對不同類型的諷刺以不同方法進行訓練,使得模型學習不同類型諷刺的特徵與關係,以偵測諷刺語句。並在最後以情感分類任務作為輔助任務,以進行多任務學習,幫助提升諷刺偵測的性能。
    我們選擇對多模態的諷刺資料集MUStARD++進行諷刺偵測,為了測試的公平性,將資料分成5個folds,做5-fold cross validation test,並以Accuracy和F1 score 作為評估指標。在Accuracy上達到了0.6575 ± 0.0260、F1則為0.6571 ± 0.0259,相較於不針對諷刺類別直接訓練約高出了0.016、0.03,且標準差也下降約0.16、0.26,表示更加穩定。結果顯示,針對不同類型制定相對應的訓練模型確實比直接對諷刺與否進行訓練有更好的表現,且加入情感分類的多任務訓練也有助於諷刺偵測。

    In recent years, there has been growing interest in sarcasm detection systems. This stems from the desire to recognize and understand the complex emotions expressed by humans. In the era of automatic response systems such as chatbots, e-commerce, and other applications becomes popular, correctly understand the intend of people's word is crucial. As misunderstandings can lead to incorrect responses.
    Currently, most research focuses on a single modality, and does not consider the different types of sarcasm identified in linguistics. Linguistics divides sarcasm into four types, with different speaker intentions and characteristics. Each type needs consideration of different information, including modalities beyond text. Therefore, we train models by different methods for different types of sarcasm, enabling the model to learn the characteristics and relationship of each type to detect sarcastic sentences. Additionally, we incorporate emotion classification as an auxiliary task in a multitask learning framework to enhance the performance of sarcasm detection.
    We choose the multimodal sarcasm dataset MUStARD++ to perform sarcasm detection.To ensure fairness in testing, the data was divided into five folds for a 5-fold cross-validation test, and we use accuracy and F1 score as evaluation metrics. The accuracy of our method reaches 0.6575 ± 0.0260 and F1 score reaches 0.6571 ± 0.0259, which are approximately 0.016 and 0.03 higher than training without distinguishing sarcasm types. The standard deviation also dropped by about 0.16 and 0.26, indicating that it is more stable. It shows that developing specific training models for different types of sarcasm performs better than directly training on sarcasm that does not consider types. Moreover, incorporating emotion classification in multitask learning further improves the sarcasm detection task.

    摘要 I Abstract III 致謝 V Table of Contents VI List of Tables IX List of Figures X Chapter 1 Introduction 1 1.1 Background 1 1.2 Motivation 3 1.3 Literature Review 4 1.3.1 Sarcasm in Linguistics 4 1.3.2 Sarcasm Detection Datasets 5 1.4 Problems 6 1.5 Brief Description of Research Methods 6 Chapter 2 Proposed Method 8 2.1 Data preparation 9 2.1.1 Data Quantity and Split 9 2.1.2 Audio Data 11 2.1.3 Text Data Augmentation 12 2.2 Model Basic Architecture 13 2.2.1 Transformer Encoder 13 2.2.2 BERT 17 2.3 Our Method 19 2.3.1 Illocutionary Model 20 2.3.2 Embedded Model 26 2.3.3 Propositional Model 28 2.3.4 Fusion 31 Chapter 3 Datasets 35 3.1 MUStARD++ 35 3.1.1 Data Sources 35 3.1.2 Sarcasm Types 36 3.1.3 Emotional Labels 38 3.2 MELD 41 Chapter 4 Experimental Setup and Result 44 4.1 Evaluation Metrics 44 4.2 Experiment Setup 46 4.2.1 Dataset 46 4.2.2 Feature Extraction 46 4.2.3 Illocutionary Model Settings 48 4.2.4 Embedded Model Settings 48 4.2.5 Propositional Model Settings 49 4.2.6 MLP Settings 49 4.2.7 Baseline 50 4.3 Results 51 4.3.1 Illocutionary Model 51 4.3.2 Embedded Model 52 4.3.3 Propositional Model 52 4.3.4 Overall Model and Ablation Study 53 Chapter 5 Conclusion and Future Work 56 Reference 57

    [1] A. Joshi, P. Bhattacharyya, and M. J. Carman, "Automatic sarcasm detection: A survey," ACM Computing Surveys (CSUR), vol. 50, no. 5, pp. 1-22, 2017.
    [2] A. Ghosh and T. Veale, "Magnets for sarcasm: Making sarcasm detection timely, contextual and very personal," in Proceedings of the 2017 conference on empirical methods in natural language processing, 2017, pp. 482-491.
    [3] R. Misra, "News headlines dataset for sarcasm detection," arXiv preprint arXiv:2212.06035, 2022.
    [4] L. Peled and R. Reichart, "Sarcasm SIGN: Interpreting sarcasm with sentiment based monolingual machine translation," arXiv preprint arXiv:1704.06836, 2017.
    [5] S. Castro, D. Hazarika, V. Pérez-Rosas, R. Zimmermann, R. Mihalcea, and S. Poria, "Towards multimodal sarcasm detection (an _obviously_ perfect paper)," arXiv preprint arXiv:1906.01815, 2019.
    [6] E. Camp, "Sarcasm, pretense, and the semantics/pragmatics distinction," Noûs, vol. 46, no. 4, pp. 587-634, 2012.
    [7] A. Ray, S. Mishra, A. Nunna, and P. Bhattacharyya, "A multimodal corpus for emotion recognition in sarcasm," arXiv preprint arXiv:2206.02119, 2022.
    [8] S. Amir, B. C. Wallace, H. Lyu, P. Carvalho, and M. J. Silva, "Modelling Context with User Embeddings for Sarcasm Detection in Social Media," in Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, 2016, pp. 167-177.
    [9] C. Van Hee, E. Lefever, and V. Hoste, "Semeval-2018 task 3: Irony detection in english tweets," in Proceedings of the 12th international workshop on semantic evaluation, 2018, pp. 39-50.
    [10] Quintilian and H. E. Butler, The Institutio Oratoria. Heinemann London, 1921.
    [11] Y. Cai, H. Cai, and X. Wan, "Multi-modal sarcasm detection in twitter with hierarchical fusion model," in Proceedings of the 57th annual meeting of the association for computational linguistics, 2019, pp. 2506-2515.
    [12] B. Liang, C. Lou, X. Li, M. Yang, L. Gui, Y. He, W. Pei, and R. Xu, "Multi-modal sarcasm detection via cross-modal graph convolutional network," in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, vol. 1: Association for Computational Linguistics, pp. 1767-1777.
    [13] H. Liu, W. Wang, and H. Li, "Towards multi-modal sarcasm detection via hierarchical congruity modeling with knowledge enhancement," arXiv preprint arXiv:2210.03501, 2022.
    [14] H. Pan, Z. Lin, P. Fu, Y. Qi, and W. Wang, "Modeling intra and inter-modality incongruity for multi-modal sarcasm detection," in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 1383-1392.
    [15] C. Wen, G. Jia, and J. Yang, "Dip: Dual incongruity perceiving network for sarcasm detection," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 2540-2550.
    [16] L. Qin, S. Huang, Q. Chen, C. Cai, Y. Zhang, B. Liang, W. Che, and R. Xu, "MMSD2. 0: Towards a Reliable Multi-modal Sarcasm Detection System," in Findings of the Association for Computational Linguistics: ACL 2023, 2023, pp. 10834-10845.
    [17] D. S. Chauhan, S. Dhanush, A. Ekbal, and P. Bhattacharyya, "Sentiment and emotion help sarcasm? A multi-task learning framework for multi-modal sarcasm, sentiment and emotion analysis," in Proceedings of the 58th annual meeting of the association for computational linguistics, 2020, pp. 4351-4360.
    [18] S. Bhosale, A. Chaudhuri, A. L. R. Williams, D. Tiwari, A. Dutta, X. Zhu, P. Bhattacharyya, and D. Kanojia, "Sarcasm in Sight and Sound: Benchmarking and Expansion to Improve Multimodal Sarcasm Detection," arXiv preprint arXiv:2310.01430, 2023.
    [19] S. Mai, Y. Sun, Y. Zeng, and H. Hu, "Excavating multimodal correlation for representation learning," Information Fusion, vol. 91, pp. 542-555, 2023.
    [20] S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea, "Meld: A multimodal multi-party dataset for emotion recognition in conversations," arXiv preprint arXiv:1810.02508, 2018.
    [21] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, "Attention is all you need," Advances in neural information processing systems, vol. 30, 2017.
    [22] R. Campos, V. Mangaravite, A. Pasquali, A. Jorge, C. Nunes, and A. Jatowt, "YAKE! Keyword extraction from single documents using multiple local features," Information Sciences, vol. 509, pp. 257-289, 2020.
    [23] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019, pp. 4171-4186.
    [24] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, and S. Anadkat, "Gpt-4 technical report," arXiv preprint arXiv:2303.08774, 2023.
    [25] M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, "BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension," in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 7871-7880.
    [26] B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, "librosa: Audio and music signal analysis in python," 2015.
    [27] K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770-778.
    [28] A. Krizhevsky, I. Sutskever, and G. E. Hinton, "Imagenet classification with deep convolutional neural networks," Advances in neural information processing systems, vol. 25, 2012.
    [29] C.-C. Hsu, S.-Y. Chen, C.-C. Kuo, T.-H. Huang, and L.-W. Ku, "EmotionLines: An Emotion Corpus of Multi-Party Conversations," in Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), 2018.
    [30] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, "Imagenet: A large-scale hierarchical image database," in 2009 IEEE conference on computer vision and pattern recognition, 2009: Ieee, pp. 248-255.

    下載圖示
    2026-08-31公開
    QR CODE