| 研究生: |
王振宇 Wang, Chen-Yu |
|---|---|
| 論文名稱: |
階層貝氏模型偵測真偽評論者之研究 Hierarchical Naïve Bayes Model for Detection of Fake Reviewers |
| 指導教授: |
李昇暾
Li, Sheng-Tun |
| 學位類別: |
碩士 Master |
| 系所名稱: |
管理學院 - 資訊管理研究所 Institute of Information Management |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 中文 |
| 論文頁數: | 49 |
| 中文關鍵詞: | 假評論者偵測 、語言探索與字詞計算 、階層貝氏模型 |
| 外文關鍵詞: | fake reviewer detection, linguistic analysis, hierarchical naïve Bayes model |
| 相關次數: | 點閱:184 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
網路的普及改變人們的生活型態,消費模式也從實體購物逐漸轉變成線上
購物,消費者在平台上相互分享與推薦產品使用經驗,這些使用經驗多少會影
響其他消費者的決定,有研究指出大約九成的消費者在選購產品時會參考已使
用過該產品的消費者評論,並且近七成的消費者信任網路評論,然而有心人士
針對這點雇用寫手在網路平台上散播假評論,想藉此來對自己增加正向評論提
高銷售,捏造負面評論來詆毀同業的聲譽,使得消費者無法正確地做出決定,因
此準確地偵測並標示假評論者成為一件重要的議題,故本研究將採用階層式分
類模型探討不同評論者所撰寫之評論對於辨別假評論者的成效。
本研究利用評論文章之語言風格特徵值來辨別假評論者,其中語言風格特
徵值包含語文探索與字詞計算(Linguistic Inquiry and Word Count, LIWC)、易讀
性(readability)、可信性(credibility)、實據性(evidentiality)與單字詞性(Part of Speech,
POS),並使用階層貝氏模型(Hierarchical Naïve Bayes Model),透過階層化同時考
慮個人與群體間之語言風格特徵,設定假評論比例門檻值以及評論數量區間,
探討各門檻值與各評論數量區間中階層貝氏模型分辨真假評論者之成效。經過
實驗後,得出階層貝氏模型相較於無考慮階層之傳統貝氏分類器(Naïve Bayes
classifier)能更有效辨別真假評論者,在門檻值為 90%且評論數量為 100 則時有
高達八成的準確率。
With the rise of online review platforms, people are used to browsing reviews on online platforms to help them make choices through other consumers' reviews before spending money. However, studies estimate that about 16% to 33% of all online reviews are fake.
Therefore, if we can come up with a method to identify them, we can greatly reduce the number of fake reviews that mislead people. We also want to investigate whether hierarchy affects the performance of the classification. Hence, we use the hierarchical naïve Bayes model to predict and identify the fake reviewers and compare it with the naïve Bayes classifier in this study. The hierarchical naïve Bayes model is an extension of the naïve Bayes classifier. We set the reviewers’ own reviews and each others’ reviews as two hierarchical levels. Considering the linguistic features of individual and group through stratification. The linguistic features of a review used in this study include Linguistic Inquiry and Word Count (LIWC), Part of Speech(POS), credibility, readability, and evidentiality. The experimental results show that the hierarchical naïve Bayes model can classify fake reviewers more effectively than naïve Bayes classifier because of the hierarchical characteristic of the data. The hierarchical naïve Bayes model’s best performance is when the threshold value is 90 and the interval of reviews is 100.
Bae, K., & Mallick, B. K. (2004). Gene selection using a two-level hierarchical Bayesian model. Bioinformatics, 20(18), 3423-3430.
Chawla, N. V., Japkowicz, N., & Kotcz, A. (2004). Special issue on learning from imbalanced data sets. ACM SIGKDD explorations newsletter, 6(1), 1-6.
Cheung, C. M.-Y., Sia, C.-L., & Kuan, K. K. Y. (2012). Is this review believable? A study of factors affecting the credibility of online consumer reviews from an ELM perspective. Journal of the Association for Information Systems, 13(8), 618-635.
Chiou, S.-Y. (2017). A trustworthy online recommendation system based on social connections in a privacy-preserving manner. Multimedia Tools and Applications, 76(7), 9319-9336.
Demichelis, F., Magni, P., Piergiorgi, P., Rubin, M. A., & Bellazzi, R. (2006). A hierarchical Naive Bayes Model for handling sample heterogeneity in classification problems: an application to tissue microarrays. BMC bioinformatics, 7(1), 514.
Fawcett, T. (2006). An introduction to ROC analysis. Pattern recognition letters, 27(8), 861-874.
Fei, G., Mukherjee, A., Liu, B., Hsu, M., Castellanos, M., & Ghosh, R. (2013). Exploiting burstiness in reviews for review spammer detection. In Proceedings of the International AAAI Conference on Web and Social Media (Vol. 7, No. 1).
Flesch, R. (1948). A new readability yardstick. Journal of applied psychology, 32(3), 221.
Ghose, A., & Ipeirotis, P. G. (2010). Estimating the helpfulness and economic impact of product reviews: Mining text and reviewer characteristics. IEEE transactions on knowledge and data engineering, 23(10), 1498-1512.
Ho, S. M., & Hancock, J. T. (2018). Computer-mediated deception: Collective language-action cues as stigmergic signals for computational intelligence. In T. Bui (Ed.), Proceedings of the 51st Hawaii International Conference on System Sciences. ScholarSpace / AIS Electronic Library (AISeL).
Jindal, N., & Liu, B. (2007). Analyzing and detecting review spam. Paper presented at the Seventh IEEE International Conference on Data Mining (ICDM 2007), Omaha, NE, United States.
Jindal, N., & Liu, B. (2008). Opinion spam and analysis. In M. A. Najork(Ed.), Proceedings of the 2008 International Conference on Web Search and Data Mining. Association for Computing Machinery.
Karklin, Y., & Lewicki, M. S. (2005). A hierarchical Bayesian model for learning nonlinear statistical regularities in nonstationary natural signals. Neural computation, 17(2), 397-423.
Katawetawaraks, C., & Wang, C. L. (2011). Online shopper behavior: Influences of online shopping decision. Asian Journal of Business Research, 1(2), 66-75.
Kauffmann, E., Peral, J., Gil, D., Ferrández, A., Sellers, R., & Mora, H. (2020). A framework for big data analytics in commercial social networks: A case study on sentiment analysis and fake review detection for marketing decision-making. Industrial Marketing Management, 90, 523-537.
Kumar, N., Venugopal, D., Qiu, L., & Kumar, S. (2018). Detecting review manipulation on online platforms with hierarchical supervised learning. Journal of Management Information Systems, 35(1), 350-380.
Lappas, T., Sabnis, G., & Valkanas, G. (2016). The impact of fake reviews on online visibility: A vulnerability assessment of the hotel industry. Information Systems Research, 27(4), 940-961.
Li, S.-T., Pham, T.-T., & Chuang, H.-C. (2019). Do reviewers’ words affect predicting their helpfulness ratings? Locating helpful reviewers by linguistics styles. Information & Management, 56(1), 28-38.
Luca, M. (2016). Reviews, reputation, and revenue: The case of Yelp.com. Harvard Business School NOM Unit Working Paper No. 12-016.
Mukherjee, A., Liu, B., Wang, J., Glance, N., & Jindal, N. (2011). Detecting group review spam. In Proceedings of the 20th international conference companion on World wide web (pp. 93-94).
Mukherjee, A., Venkataraman, V., Liu, B., & Glance, N. S. (2013). What yelp fake review filter might be doing? Paper presented at the Seventh International AAAI Conference on Weblogs and Social Media, Cambridge, MA, United States.
Ott, M., Choi, Y., Cardie, C., & Hancock, J. T. (2011). Finding deceptive opinion spam by any stretch of the imagination. In D. Lin (Ed.), Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human language technologies-volume 1. Association for Computational Linguistics.
Pérez-Rosas, V., Kleinberg, B., Lefevre, A., & Mihalcea, R. (2018). Automatic detection of fake news. In E. M. Bender, L. Derczynski, & P. Isabelle (Eds.), Proceedings of the 27th International Conference on Computational Linguistics. Association for Computational Linguistics.
Pennebaker, J. W., Boyd, R. L., Jordan, K., & Blackburn, K. (2015). The development and psychometric properties of LIWC2015. Austin, TX: University of Texas at Austin. Retrieved from https://repositories.lib.utexas.edu/bitstream/handle/2152/31333/
LIWC2015_LanguageManual.pdf
Pikulski, J. J. (2002). Readability. Boston: Houghton Mifflin. Retrieved from http://www.eduplace.com/state/author/pikulski.pdf
Reichheld, F. F. (2003). The one number you need to grow. Harvard business review. Retrieved from https://hbr.org/2003/12/the-one-number-you-need-to-grow
Rish, I. (2001). An empirical study of the naive Bayes classifier. In IJCAI 2001 workshop on empirical methods in artificial intelligence (Vol. 3, No. 22, pp. 41-46).
Su, Q., Huang, C.-R., & Chen, H. K. (2010). Evidentiality for text trustworthiness detection. In F. Xia, W. Lewis, & L. Levin (Eds.), Proceedings of the 2010 Workshop on NLP and Linguistics: Finding the Common Ground. Association for Computational Linguistics.
Wang, F., & Karimi, S. (2018). Linguistic Style and Online Review Helpfulness. Paper presented at the Thirty Eighth International Conference on Information Systems, South Korea.
Wang, G., Xie, S., Liu, B., & Yu, P. S. (2012). Identify online store review spammers via social review graph. ACM Transactions on Intelligent Systems and Technology (TIST), 3(4), 1-21.
Wu, Y., Ngai, E. W. T., Wu, P., & Wu, C. (2020). Fake online reviews: Literature review, synthesis, and directions for future research. Decision Support Systems, 132(5), 113280.