| 研究生: |
吳佳穎 Wu, Chia-Ying |
|---|---|
| 論文名稱: |
應用主題模型之論據子議題發掘與立場偵測方法 Stance Detection with Argument Facet Discovery based on Neural Topic Model |
| 指導教授: |
王惠嘉
Wang, Hei-Chia |
| 學位類別: |
碩士 Master |
| 系所名稱: |
管理學院 - 資訊管理研究所 Institute of Information Management |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 中文 |
| 論文頁數: | 59 |
| 中文關鍵詞: | 論據探勘 、類神經主題模型 、立場偵測 |
| 外文關鍵詞: | Argument Mining, Neural Topic Model, Stance Detection |
| 相關次數: | 點閱:140 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
辯論網站為用於討論具有爭議性議題的開放平台,根據使用者提出的論述分為支持及反對兩方,討論的議題涉及經濟、政治、科技、文化、環境等多領域。隨著全球網路使用者的人數所佔比例逐年提升,使用者生成內容數量亦快速成長,熱門議題的平均論據數量高達數千則,對於決策者或是不熟悉議題者想要瞭解使用者的想法時,需要逐一瀏覽耗費大量時間才能從中得到主要論點的總結。隨著自然語言處理技術的發展,學者陸續提出自動識別和提取論點結構的論據探勘方法,其中立場偵測任務為根據論述判斷該使用者的立場,然而此類研究追求分類正確率,結果所能提供的資訊有限,對於不熟悉該議題的人來說,除了瞭解論據正反方組成之外,造成兩派之間意見無法達成共識的核心更為重要。
在辯論討論中存在造成立場兩極的子議題,使用者的論點會從不同子議題切入,相同立場的使用者的所討論的子議題也不盡相同。提取文本潛在主題的方法根據推論過程採用架構分為機率模型及類神經主題模型,後者相較於前者改善了推導計算上的成本且能夠提取一致性更高的主題。因此,本研究提出在立場偵測的基礎上加入子議題發掘的兩階段模型,首先以變分自編碼器架構的類神經主題模型取得各個論據文本中潛在子議題的分布,接著在立場偵測階段,使用循環神經網路學習論據特徵,考量不同子議題中論據立場的闡述方式,除了輸入論據文本之外,也將納入該論據的子議題分布特徵增強模型辨別立場的能力。經由實驗結果指出,本研究所提出的子議題模型相較於過去相關文獻方法能夠取得連貫性更高的子議題,且觀察發現代表字確實反應出議題下不同面向討論的關鍵字,而在立場偵測模型達到近七成的分類效果,顯示子議題特徵對於模型辨別立場的有效性。
Debate website is a platform for controversial issue discussing. Users express their stance and understand the ideas of users from different stance through the platform. As global Internet users has increased, the amount of user-generated content has grown rapidly. The quantity of arguments on popular issues is up to thousands. For those who are unfamiliar with issues, it is time-consuming to browse through the discussion to get a summary of the main viewpoints.
With the development of natural language processing, scholars have proposed methods for automatically identifying argument structures. Stance detection is to judge the user’s stance based on the argument. In addition to the stance, the facets that the two factions cannot reach a consensus is more important. The method of extracting potential topics is divided into probability models and neural topic models according to the inference process. The latter improves computational cost of derivation and can extract more consistent topics than the former.
This research proposes a two-stage model. First, a neural topic model based on variational autoencoder is used to obtain the potential subtopic distribution. At stance detection stage, recurrent neural network is used to learn the characteristics of the argument. Facet distribution of the argument is also included as input to enhance the model's ability. The experimental results show that the model can obtain more coherent subtopics compared to previous studies. The stance detection model achieves accuracy of nearly 70%, showing the effectiveness of the subtopic features.
Abbott, R., Ecker, B., Anand, P., & Walker, M. (2016). Internet Argument Corpus 2.0: An Sql Schema for Dialogic Social Media and the Corpora to Go with It. Paper presented at the Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), Portorož, Slovenia.
Augenstein, I., Rocktäschel, T., Vlachos, A., & Bontcheva, K. (2016). Stance Detection with Bidirectional Conditional Encoding. Paper presented at the Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Austin, Texas, USA.
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural Machine Translation by Jointly Learning to Align and Translate. Paper presented at the arXiv preprint arXiv:1409.0473.
Bar-Haim, R., Bhattacharya, I., Dinuzzo, F., Saha, A., & Slonim, N. (2017). Stance Classification of Context-Dependent Claims. Paper presented at the Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, Valencia, Spain.
Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent Dirichlet Allocation. Journal of machine Learning research, 3(Jan), 993-1022.
Chang, J., Boyd-Graber, J., Wang, C., Gerrish, S., & Blei, D. M. (2009). Reading Tea Leaves: How Humans Interpret Topic Models. Paper presented at the Neural information processing systems, Vancouver, Canada.
Chen, W.-F., & Ku, L.-W. (2016). Utcnn: A Deep Learning Model of Stance Classificationon on Social Media Text. Paper presented at the arXiv preprint arXiv:1611.03599.
Chowanda, A. D., Sanyoto, A. R., Suhartono, D., & Setiadi, C. J. (2017). Automatic Debate Text Summarization in Online Debate Forum. Procedia computer science, 116, 11-19.
Chung, J., Gulcehre, C., Cho, K., & Bengio, Y. (2014). Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. Paper presented at the arXiv preprint arXiv:1412.3555.
Deerwester, S., Dumais, S. T., Furnas, G. W., Landauer, T. K., & Harshman, R. (1990). Indexing by Latent Semantic Analysis. Journal of the American society for information science, 41(6), 391-407.
Ding, R., Nallapati, R., & Xiang, B. (2018). Coherence-Aware Neural Topic Modeling. Paper presented at the arXiv preprint arXiv:1809.02687.
Du, J., Xu, R., He, Y., & Gui, L. (2017). Stance Classification with Target-Specific Neural Attention Networks. Paper presented at the International Joint Conferences on Artificial Intelligence, Melbourne, Australia.
Ebrahimi, J., Dou, D., & Lowd, D. (2016). A Joint Sentiment-Target-Stance Model for Stance Classification in Tweets. Paper presented at the Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics, Osaka, Japan.
Feng, J., Rao, Y., Xie, H., Wang, F. L., & Li, Q. (2019). User Group Based Emotion Detection and Topic Discovery over Short Text. World Wide Web, 1-35.
Feng, J., Zhang, Z., Ding, C., Rao, Y., & Xie, H. (2020). Context Reinforced Neural Topic Modeling over Short Texts. Paper presented at the arXiv preprint arXiv:2008.04545.
Ghannay, S., Favre, B., Esteve, Y., & Camelin, N. (2016). Word Embedding Evaluation and Combination. Paper presented at the Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), Portorož, Slovenia.
Gleize, M., Shnarch, E., Choshen, L., Dankin, L., Moshkowich, G., Aharonov, R., & Slonim, N. (2019). Are You Convinced? Choosing the More Convincing Evidence with a Siamese Network. Paper presented at the Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy.
Griffiths, T. L., & Steyvers, M. (2004). Finding Scientific Topics. Proceedings of the National academy of Sciences, 101(suppl 1), 5228-5235.
He, R., Lee, W. S., Ng, H. T., & Dahlmeier, D. (2017). An Unsupervised Neural Attention Model for Aspect Extraction. Paper presented at the Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vancouver, Canada.
Hinton, G. E., & Salakhutdinov, R. R. (2006). Reducing the Dimensionality of Data with Neural Networks. science, 313(5786), 504-507.
Hochreiter, S., & Schmidhuber, J. (1997). Long Short-Term Memory. Neural computation, 9(8), 1735-1780.
Hofmann, T. (1999). Probabilistic Latent Semantic Indexing. Paper presented at the Proceedings of the 22nd annual international ACM SIGIR conference on Research and development in information retrieval, Berkeley, California, USA.
Jin, O., Liu, N. N., Zhao, K., Yu, Y., & Yang, Q. (2011). Transferring Topical Knowledge from Auxiliary Long Texts for Short Text Clustering. Paper presented at the Proceedings of the 20th ACM international conference on Information and knowledge management, Glasgow Scotland, UK.
Kemp, S. (2020). Digital 2020: Global Digital Overview. Retrieved from https://datareportal.com/reports/digital-2020-global-digital-overview
Kingma, D. P., & Welling, M. (2013). Auto-Encoding Variational Bayes. Paper presented at the arXiv preprint arXiv:1312.6114.
Lawrence, J., & Reed, C. (2020). Argument Mining: A Survey. Computational Linguistics, 45(4), 765-818.
Levy, R., Bogin, B., Gretz, S., Aharonov, R., & Slonim, N. (2018). Towards an Argumentative Content Search Engine Using Weak Supervision. Paper presented at the Proceedings of the 27th International Conference on Computational Linguistics, Santa Fe, New Mexico, USA.
Li, C., Porco, A., & Goldwasser, D. (2018). Structured Representation Learning for Online Debate Stance Prediction. Paper presented at the Proceedings of the 27th International Conference on Computational Linguistics, Santa Fe, New Mexico, USA.
Liu, L., Huang, H., Gao, Y., Zhang, Y., & Wei, X. (2019). Neural Variational Correlated Topic Modeling. Paper presented at the The World Wide Web Conference, San Francisco, California.
McClelland, J. L., & Rumelhart, D. E. (1987). Parallel Distributed Processing (Vol. 2): MIT press Cambridge, MA:.
McCulloch, W. S., & Pitts, W. (1943). A Logical Calculus of the Ideas Immanent in Nervous Activity. The bulletin of mathematical biophysics, 5(4), 115-133.
Miao, Y., Grefenstette, E., & Blunsom, P. (2017). Discovering Discrete Latent Topics with Neural Variational Inference. Paper presented at the arXiv preprint arXiv:1706.00359.
Miao, Y., Yu, L., & Blunsom, P. (2016). Neural Variational Inference for Text Processing. Paper presented at the International conference on machine learning, New York, USA.
Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. Paper presented at the arXiv preprint arXiv:1301.3781.
Misra, A., Anand, P., Tree, J. E. F., & Walker, M. (2015). Using Summarization to Discover Argument Facets in Online Idealogical Dialog. Paper presented at the Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Denver, Colorado, USA.
Misra, A., Ecker, B., & Walker, M. A. (2017). Measuring the Similarity of Sentential Arguments in Dialog. Paper presented at the arXiv preprint arXiv:1709.01887.
Mohammad, S., Kiritchenko, S., Sobhani, P., Zhu, X., & Cherry, C. (2016). Semeval-2016 Task 6: Detecting Stance in Tweets. Paper presented at the Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016), San Diego, California, USA.
Newman, D., Lau, J. H., Grieser, K., & Baldwin, T. (2010). Automatic Evaluation of Topic Coherence. Paper presented at the Human language technologies: The 2010 annual conference of the North American chapter of the association for computational linguistics, Los Angeles, California.
Paul, M., Zhai, C., & Girju, R. (2010). Summarizing Contrastive Viewpoints in Opinionated Text. Paper presented at the Proceedings of the 2010 conference on empirical methods in natural language processing, Massachusetts, USA.
Pennington, J., Socher, R., & Manning, C. D. (2014). Glove: Global Vectors for Word Representation. Paper presented at the Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), Doha, Qatar.
Ranade, S., Gupta, J., Varma, V., & Mamidi, R. (2013). Online Debate Summarization Using Topic Directed Sentiment Analysis. Paper presented at the Proceedings of the Second International Workshop on Issues of Sentiment Discovery and Opinion Mining, Chicago, Illinois, USA.
Rezende, D. J., Mohamed, S., & Wierstra, D. (2014). Stochastic Backpropagation and Approximate Inference in Deep Generative Models. Paper presented at the International Conference on Machine Learning, Beijing, China.
Sasaki, A., Hanawa, K., Okazaki, N., & Inui, K. (2018). Predicting Stances from Social Media Posts Using Factorization Machines. Paper presented at the Proceedings of the 27th International Conference on Computational Linguistics, Santa Fe, New Mexico, USA.
Seok, M., Song, H.-J., Park, C.-Y., Kim, J.-D., & Kim, Y.-s. (2016). Named Entity Recognition Using Word Embedding as a Feature. International Journal of Software Engineering and Its Applications, 10(2), 93-104.
Sethuraman, J. (1994). A Constructive Definition of Dirichlet Priors. Statistica sinica, 639-650.
Sobhani, P., Mohammad, S., & Kiritchenko, S. (2016). Detecting Stance in Tweets and Analyzing Its Interaction with Sentiment. Paper presented at the Proceedings of the fifth joint conference on lexical and computational semantics, Berlin, Germany.
Srivastava, A., & Sutton, C. (2017). Autoencoding Variational Inference for Topic Models. Paper presented at the arXiv preprint arXiv:1703.01488.
Sun, Q., Wang, Z., Zhu, Q., & Zhou, G. (2016). Exploring Various Linguistic Features for Stance Detection. In Natural Language Understanding and Intelligent Applications (pp. 840-847): Springer.
Tarwani, K. M., & Edem, S. (2017). Survey on Recurrent Neural Network in Natural Language Processing. Int. J. Eng. Trends Technol, 48, 301-304.
Thonet, T., Cabanac, G., Boughanem, M., & Pinel-Sauvagnat, K. (2016). Vodum: A Topic Model Unifying Viewpoint, Topic and Opinion Discovery. Paper presented at the European conference on information retrieval, Padua, Italy.
Trabelsi, A., & Zaiane, O. R. (2018). Unsupervised Model for Topic Viewpoint Discovery in Online Debates Leveraging Author Interactions. Paper presented at the Proceedings of the International AAAI Conference on Web and Social Media, Stanford, California.
Trabelsi, A., & Zaïane, O. R. (2019). Phaitv: A Phrase Author Interaction Topic Viewpoint Model for the Summarization of Reasons Expressed by Polarized Stances. Paper presented at the Proceedings of the International AAAI Conference on Web and Social Media, Munich, Germany.
Vilares, D., & He, Y. (2017). Detecting Perspectives in Political Debates. Paper presented at the Proceedings of the 2017 conference on empirical methods in natural language processing, Copenhagen, Denmark.
Wang, Z., Sun, Q., Li, S., Zhu, Q., & Zhou, G. (2020). Neural Stance Detection with Hierarchical Linguistic Representations. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28, 635-645.
Wei, W., Zhang, X., Liu, X., Chen, W., & Wang, T. (2016). Pkudblab at Semeval-2016 Task 6: A Specific Convolutional Neural Network System for Effective Stance Detection. Paper presented at the Proceedings of the 10th international workshop on semantic evaluation (SemEval-2016), San Diego, California, USA.
Xu, C., Paris, C., Nepal, S., & Sparks, R. (2018). Cross-Target Stance Classification with Self-Attention Networks. Paper presented at the Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Melbourne, Australia.
Zhao, W. X., Jiang, J., Weng, J., He, J., Lim, E.-P., Yan, H., & Li, X. (2011). Comparing Twitter and Traditional Media Using Topic Models. Paper presented at the European conference on information retrieval, Dublin, Ireland.
Zhu, Q., Feng, Z., & Li, X. (2018). Graphbtm: Graph Enhanced Autoencoded Variational Inference for Biterm Topic Model. Paper presented at the Proceedings of the 2018 conference on empirical methods in natural language processing, Brussels, Belgium.