| 研究生: |
李文丁 Li, Wen-Ding |
|---|---|
| 論文名稱: |
社群媒體分析為基之企業信用評等方法研究 On Social Media Analytics-based Corporate Credit Rating Method |
| 指導教授: |
陳裕民
Chen, Yuh-Min |
| 共同指導: |
陳育仁
Chen, Yuh-Jen |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 製造資訊與系統研究所 Institute of Manufacturing Information and Systems |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 中文 |
| 論文頁數: | 54 |
| 中文關鍵詞: | 企業信用評等 、信用評等預測 、社群媒體分析 、大數據 |
| 外文關鍵詞: | credit rating, credit rating prediction, social media, big data |
| 相關次數: | 點閱:168 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
企業信用評等因2007年雷曼兄弟引發的金融危機而受到重視。傳統上,金融機構大多以企業財務指標或公司治理指標,進行企業信用評等。隨著網際網路的發達以及社群媒體的普及,在社群媒體也提供金融機構另一種客觀的資料來源,能用於衡量企業信用評等。因此,有效地從大量且雜亂的網路社群媒體資料中,分析與預測企業信用評等也成為金融科技(FinTech)浪潮下,金融機構風險管理部門之課題。
針對金融機構對於企業信用評等之需求以及社群媒體內容的可用性,本研究主要目的在於設計一個社群媒體分析為基之企業信用評等預測方法,並開發其實現技術,以協助金融機構有效進行企業風險評估與控管。
本方法首先針對網路中社群平台的大數據,擷取企業之相關評論,接著分析社群平台使用者對特定企業的觀點與評價,並透過預測模型獲取該企業於最新一季之信用風險評等。
針對上述目的,本研究主要研究項目包括:(i)網路社群媒體大數據為基之企業信用評等預測方法設計,(ii)網路社群媒體大數據為基之企業信用評等預測方法發展,以及(iii)網路社群媒體大數據為基之企業信用評等預測系統實作與評估。
為驗證方法之有效性,本研究經由實驗,先找出以企業輿情指標訓練模型的最佳組合,再找出模型最佳的超參數組合,最終得到模型建構的最佳配置,在經過10折的交叉驗證後,模型的準確度達到81.07%。接著將該模型套用於本研究所提之方法中,並實際用於預測,最終得出準確度為78.4%,相較於模型建構時測試的準確度81.07%,兩者之間僅差距2.67%。故本研究所提以社群媒體分析為基之企業信用評等預測方法可行且有效。
The corporate credit rating has received attention due to the financial crisis triggered by Lehman Brothers in 2007. Traditionally, financial institutions mostly use corporate financial indicators or corporate governance indicators to conduct corporate credit ratings. With the development of the Internet and the popularization of social media, social media also provides another objective source of information for financial institutions, which can be used to measure corporate credit ratings. Therefore, effectively analyzing and predicting corporate credit ratings from the large and messy online social media data has become a task for the risk management department of financial institutions under the wave of FinTech.
In response to the needs of financial institutions for corporate credit ratings and the availability of social media content, the main purpose of this research is to design a social media analysis-based corporate credit rating prediction method and develop its implementation technology to assist financial institutions effectively conduct corporate risk assessment and control.
This method first focuses on the big data of the social network platform on the Internet, extracts the relevant comments of the corporate, then analyzes the opinions and evaluations of the users of the social platform on a specific corporate, and obtains the corporate’s credit risk evaluation in the latest quarter through a predictive model. Etc.
In view of the above-mentioned purposes, the main research projects of this research include: (i) the design of prediction methods for a corporate credit rating based on online social media big data, (ii) the prediction of corporate credit rating based on online social media big data Method development, and (iii) the implementation and evaluation of a corporate credit rating prediction system based on big data in online social media.
In order to verify the effectiveness of the method, this research first finds the best combination of training models with corporate public opinion indicators through experiments, and then finds the best combination of hyperparameters for the model, and finally obtains the best configuration for model construction. After cross-validation, the accuracy of the model reached 81.07%. Then apply the model to the method proposed in this study and actually use it for prediction. The final accuracy is 78.4%. Compared with the accuracy of 81.07% tested during model construction, the difference between the two is only 2.67%. . Therefore, the method of predicting corporate credit rating based on social media analysis proposed in this research is feasible and effective.
[1] Abdollahi, M., T. Khaleghi, and K. Yang, An integrated feature learning approach using deep learning for travel time prediction. Expert Systems with Applications, 2020. 139: p. 112864.
[2] Belgiu, M. and L. Drăguţ, Random forest in remote sensing: A review of applications and future directions. ISPRS journal of photogrammetry and remote sensing, 2016. 114: p. 24-31.
[3] Berrar, D., Bayes’ theorem and naive Bayes classifier. Encyclopedia of Bioinformatics and Computational Biology: ABC of Bioinformatics; Elsevier Science Publisher: Amsterdam, The Netherlands, 2018: p. 403-412.
[4] Blei, D.M., A.Y. Ng, and M.I. Jordan, Latent dirichlet allocation. the Journal of machine Learning research, 2003. 3: p. 993-1022.
[5] Breiman, L., Random forests. Machine learning, 2001. 45(1): p. 5-32.
[6] Chawla, N.V., et al., SMOTE: synthetic minority over-sampling technique. Journal of artificial intelligence research, 2002. 16: p. 321-357.
[7] Chen, T. and C. Guestrin. Xgboost: A scalable tree boosting system. in Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining. 2016.
[8] Daniel, C., J. Hančlová, and H. el Woujoud Bousselmi, Corporate rating forecasting using Artificial Intelligence statistical techniques. Investment Management & Financial Innovations, 2019. 16(2): p. 295.
[9] Darling, W.M. A theoretical and practical implementation tutorial on topic modeling and gibbs sampling. in Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies. 2011.
[10] Delen, D. and H. Demirkan, Data, information and analytics as services. 2013, Elsevier.
[11] Fei, W., et al., Credit risk evaluation based on social media. Procedia Computer Science, 2015. 55: p. 725-731.
[12] Gül, S., Ö. Kabak, and I. Topcu, A multiple criteria credit rating approach utilizing social media data. Data & Knowledge Engineering, 2018. 116: p. 80-99.
[13] Granik, M. and V. Mesyura. Fake news detection using naive Bayes classifier. in 2017 IEEE first Ukraine conference on electrical and computer engineering (UKRCON). 2017. IEEE.
[14] Guo, G., et al. KNN model-based approach in classification. in OTM Confederated International Conferences" On the Move to Meaningful Internet Systems". 2003. Springer.
[15] Hagen, L., Content analysis of e-petitions with topic modeling: How to train and evaluate LDA models? Information Processing & Management, 2018. 54(6): p. 1292-1307.
[16] Hajjem, M. and C. Latiri, Combining IR and LDA topic modeling for filtering microblogs. Procedia Computer Science, 2017. 112: p. 761-770.
[17] Harrou, F., A. Zeroual, and Y. Sun, Traffic congestion monitoring using an improved kNN strategy. Measurement, 2020. 156: p. 107534.
[18] Hsu, F.-J., M.-Y. Chen, and Y.-C. Chen, The human-like intelligence with bio-inspired computing approach for credit ratings prediction. Neurocomputing, 2018. 279: p. 11-18.
[19] Hutto, C. and E. Gilbert. Vader: A parsimonious rule-based model for sentiment analysis of social media text. in Proceedings of the International AAAI Conference on Web and Social Media. 2014.
[20] Joachims, T., Making large scale SVM learning practical (No. 1998, 28). 1999, Technical Report, SFB 475: Komplexitätsreduktion in Multivariaten ….
[21] Jones, S., D. Johnstone, and R. Wilson, An empirical evaluation of the performance of binary classifiers in the prediction of credit ratings changes. Journal of Banking & Finance, 2015. 56: p. 72-85.
[22] Kalogirou, S.A., Applications of artificial neural-networks for energy systems. Applied energy, 2000. 67(1-2): p. 17-35.
[23] Karminsky, A.M. and E. Khromova, Extended modeling of banks’ credit ratings. Procedia Computer Science, 2016. 91: p. 201-210.
[24] Kavaklioglu, K., et al., Modeling and prediction of Turkey’s electricity consumption using artificial neural networks. Energy Conversion and Management, 2009. 50(11): p. 2719-2727.
[25] Khemakhem, S. and Y. Boujelbene, Predicting credit risk on the basis of financial and non-financial variables and data mining. Review of Accounting and Finance, 2018.
[26] Kim, K.-j. and H. Ahn, A corporate credit rating model using multi-class support vector machines with an ordinal pairwise partitioning approach. Computers & Operations Research, 2012. 39(8): p. 1800-1811.
[27] Ku, L.W. and H.H. Chen, Mining opinions from the Web: Beyond relevance retrieval. Journal of the American Society for Information Science and Technology, 2007. 58(12): p. 1838-1850.
[28] Lee, I., Social media analytics for enterprises: Typology, methods, and processes. Business Horizons, 2018. 61(2): p. 199-210.
[29] Lee, S., et al., Determination and application of the weights for landslide susceptibility mapping using an artificial neural network. Engineering Geology, 2004. 71(3-4): p. 289-302.
[30] Li, G., et al., Using lda model to quantify and visualize textual financial stability report. Procedia computer science, 2017. 122: p. 370-376.
[31] Lin, Y., H. Guo, and J. Hu. An SVM-based approach for stock market trend prediction. in The 2013 international joint conference on neural networks (IJCNN). 2013. IEEE.
[32] Liu, B., et al. Scalable sentiment classification for big data analysis using naive bayes classifier. in 2013 IEEE international conference on big data. 2013. IEEE.
[33] Liu, H., et al., Combining user preferences and user opinions for accurate recommendation. Electronic Commerce Research and Applications, 2013. 12(1): p. 14-23.
[34] Liu, H. and Z. Long, An improved deep learning model for predicting stock market price time series. Digital Signal Processing, 2020. 102: p. 102741.
[35] Loader, B.D., Social movements and new media. Sociology Compass, 2008. 2(6): p. 1920-1933.
[36] Ma, W.-Y. and K.-J. Chen. Introduction to CKIP Chinese word segmentation system for the first international Chinese word segmentation bakeoff. in Proceedings of the second SIGHAN workshop on Chinese language processing. 2003.
[37] Mayer-Schönberger, V. and K. Cukier, Big data: A revolution that will transform how we live, work, and think. 2013: Houghton Mifflin Harcourt.
[38] McInnes, L., J. Healy, and J. Melville, Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018.
[39] Moldagulova, A. and R.B. Sulaiman. Using KNN algorithm for classification of textual documents. in 2017 8th International Conference on Information Technology (ICIT). 2017. IEEE.
[40] Ogunleye, A. and Q.-G. Wang, XGBoost model for chronic kidney disease diagnosis. IEEE/ACM transactions on computational biology and bioinformatics, 2019. 17(6): p. 2131-2140.
[41] Park, S.B., et al., Visualizing theme park visitors’ emotions using social media analytics and geospatial analytics. Tourism Management, 2020. 80: p. 104127.
[42] Pathak, A.R., et al., Analysis of Techniques for Rumor Detection in Social Media. Procedia Computer Science, 2020. 167: p. 2286-2296.
[43] Reynolds, W.N., et al. Social media and social reality. in 2010 IEEE International Conference on Intelligence and Security Informatics. 2010. IEEE.
[44] Rish, I. An empirical study of the naive Bayes classifier. in IJCAI 2001 workshop on empirical methods in artificial intelligence. 2001.
[45] Rumelhart, D.E., G.E. Hinton, and R.J. Williams, Learning representations by back-propagating errors. nature, 1986. 323(6088): p. 533-536.
[46] Sauter, D.A., et al., Cross-cultural recognition of basic emotions through nonverbal emotional vocalizations. Proceedings of the National Academy of Sciences, 2010. 107(6): p. 2408-2412.
[47] Simonofski, A., J. Fink, and C. Burnay, Supporting policy-making with social media and e-participation platforms data: A policy analytics framework. Government Information Quarterly, 2021: p. 101590.
[48] Sproat, R. and C. Shih, A statistical method for finding word boundaries in Chinese text. Computer Processing of Chinese and Oriental Languages, 1990. 4(4): p. 336-351.
[49] Stieglitz, S., et al., Social media analytics–Challenges in topic discovery, data collection, and data preparation. International journal of information management, 2018. 39: p. 156-168.
[50] Stjepan, O., O. Dijana, and O. Goran, Hybrid system with genetic algorithm and artificial neural networks and its application to retail credit risk assessment. Expert systems with applications, 2012. 39(16): p. 12605-12617.
[51] Sun, J., Jieba chinese word segmentation tool. Accessed: Jun, 2012. 25: p. 2018.
[52] Tang, H., S. Tan, and X. Cheng, A survey on sentiment detection of reviews. Expert Systems with Applications, 2009. 36(7): p. 10760-10773.
[53] Tian, X., et al., Latency critical big data computing in finance. The Journal of Finance and Data Science, 2015. 1(1): p. 33-41.
[54] Turney, P.D., Thumbs up or thumbs down? Semantic orientation applied to unsupervised classification of reviews. arXiv preprint cs/0212032, 2002.
[55] Vishwanathan, S. and M.N. Murty. SSVM: a simple SVM algorithm. in Proceedings of the 2002 International Joint Conference on Neural Networks. IJCNN'02 (Cat. No. 02CH37290). 2002. IEEE.
[56] Wallis, M., K. Kumar, and A. Gepp, Credit rating forecasting using machine learning techniques, in Managerial Perspectives on Intelligent Big Data Analytics. 2019, IGI Global. p. 180-198.
[57] Wang, G. and J. Ma, Study of corporate credit risk prediction based on integrating boosting and random subspace. Expert Systems with Applications, 2011. 38(11): p. 13871-13878.
[58] Xing, W. and Y. Bei, Medical health big data classification based on KNN classification algorithm. IEEE Access, 2019. 8: p. 28808-28819.
[59] Yeh, C.-C., F. Lin, and C.-Y. Hsu, A hybrid KMV model, random forests and rough set theory approach for credit rating. Knowledge-Based Systems, 2012. 33: p. 166-172.
[60] You, Z., et al., A decision-making framework for precision marketing. Expert Systems with Applications, 2015. 42(7): p. 3357-3367.
[61] Yuan, Z., et al., Unsupervised multi-granular Chinese word segmentation and term discovery via graph partition. Journal of Biomedical Informatics, 2020. 110: p. 103542.
[62] Zhang, M.-L. and Z.-H. Zhou, ML-KNN: A lazy learning approach to multi-label learning. Pattern recognition, 2007. 40(7): p. 2038-2048.
[63] 王宏元(2018)。應用網路社群媒體大數據分析建立企業貸款信用風險評估機制。 高雄科技大學會計資訊研究所學位論文,1-73。
[64] 吳俊翰(2017)。網路資料為基之產品演化歷程挖掘與預測方法研發。成功大學製造資訊與系統研究所學位論文,1-51。
[65] 林美鳳、金成隆、林良楓(2009)。股權結構、會計保守性與信用評等關係之研究。臺大管理論叢,20(1),289-329。
[66] 洪健哲(2019)。社群媒體為基之市場區隔趨勢預測方法與技術研發。成功大學製造資訊與系統研究所學位論文,1-90。
[67] 高聖傑(2014)。詞性組合輔助之中文網路口碑評價分析技術研發。成功大學製造資訊與系統研究所學位論文,1-59。
[68] 游和正、黃挺豪、陳信希(2012)。領域相關詞彙極性分析及文件情緒分類之研究。中文計算語言學期刊,17(4),33-47。