| 研究生: |
辛添佑 SHIN, Tian-You |
|---|---|
| 論文名稱: |
在提升法中導入用簡易貝氏分類器來隨機生成基本模型之研究 A Study on Introducing Randomly-Generated Base Models in Boosting Approach for Naïve Bayesian Classifiers |
| 指導教授: |
翁慈宗
Wong, Tzu-Tsung |
| 學位類別: |
碩士 Master |
| 系所名稱: |
管理學院 - 工業與資訊管理學系 Department of Industrial and Information Management |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 70 |
| 中文關鍵詞: | 集成學習 、提升法 、隨機生成基本模型方法 、簡易貝氏分類器 |
| 外文關鍵詞: | Ensemble Learning, Boosting, Randomly Generated Base Models, Naïve Bayesian Classifier, Particle Swarm Optimization |
| 相關次數: | 點閱:2 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
傳統集成學習的研究多著重於解決基本模型間相依性過高的問題,相依性低的要求之所以產生,係因基本模型之間皆使用相同的訓練資料集所導致,即使透過訓練資料集進行隨機抽樣亦或是加權調整,仍無法有效改善。
故後續研究跳脫出訓練資料集的框架,不依賴訓練資料集進行學習,而是透過隨機的方式生成基本模型,並且在與袋裝法結合的實驗中,展現出良好的效能以及發展性,然而現有研究普遍缺乏將此方法與提升法結合的探討,主要原因為提升法的基本模型生成具有順序性,使其難以與隨機生成基本模型方法進行整合,故本研究設計兩種提升法結合隨機生成基本模型方法,使得在提升法架構下也可以運用隨機生成的基本模型,並以簡易貝氏分類器作為基本模型,由於其運算效率佳,且在隨機生成基本模型下,實作較為容易。
本研究旨在探討將提升法與隨機生成基本模型方法相結合後,是否能有效提升整體預測效能,並進一步與過往袋裝法混合隨機生成基本模型方法比較。研究結果顯示出將隨機生成基本模型方法與提升法結合後,平均正確率有上升,並且發現在提升法比起袋裝法有更高分類正確率的資料集上,繼續使用提升法結合隨機生成基本模型方法可以保持優勢。
This study explores whether randomly generated base models can be combined with boosting to improve classification performance. Traditional ensemble learning methods often suffer from high dependency among base models,because their base models are usually generated through similar learning processes. Previous studies have demonstrated that integrating randomly generated base models into bagging frameworks can effectively improve classification accuracy. However, little research has investigated their application in boosting. Therefore, this study proposes several approaches to combine randomly generated base models with boosting and examines their effectiveness.
何政賢,(2024) 以粒子群最佳化方法優化應用於二類別資料之隨機集成演算法。國立成功大學資訊管理研究所碩士班碩士論文
黃中立,(2023) 以簡易貝氏分類器隨機生成基本模型之集成方法。國立成功大學資訊管理研究所碩士班碩士論文
黃愉兒,(2025) 抽樣方法結合隨機生成法應用於分類不平衡資料集。國立成功大學資訊管理研究所碩士班碩士論文
蔡哲倫,(2024) 適用於多類別資料的含隨機生成模型之集成嵌套二分法。國立成功大學資訊管理研究所碩士班碩士論文
Breiman,L.(1996) Bagging predictors.Machine Learning,24(2),123–140.
Breiman,L.(2001) Random forests.Machine Learning,45(1),5–32.
Budholiya,K.,Shrivastava,S.K.& Sharma,V.(2022) An optimized XGBoost based diagnostic system for effective prediction of heart disease. Journal of King Saud University – Computer and Information Sciences,34,4514–4523.
Chaudhary, A., Thakur, R., Kolhe, S., & Kamal, R. (2020). A particle swarm optimization based ensemble for vegetable crop disease recognition. Computers and Electronics in Agriculture, 178, 105747.
Chen,T.& Guestrin,C.(2016) XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD ’16), 785–794.
Dietterich,T.G.(2000) An experimental comparison of three methods for constructing ensembles of decision trees: Bagging, boosting, and randomization.Machine Learning,40,139–157.
Dietterich.T.G (2000) Ensemble Methods in Machine Learning. Proceedings of the 1st International Workshop on Multiple Classifier Systems (pp.1-15).Springer.
Faridah,N.,Dewi,C.& Soebroto,A.A.(2021) Hybrid of AdaBoost algorithm and Naïve Bayes classifier on selection of contraception methods. Journal of Environmental Engineering and Sustainable Technology, 8(2),107–116.
Gedela,B.& Karthikeyan,P.R.(2022, February) Credit card fraud detection using AdaBoost algorithm in comparison with various machine learning algorithms to measure accuracy, sensitivity, specificity, precision and F-score. Paper presented at 2022 International Conference on Business Analytics for Technology and Security (ICBATS), IEEE.
Hashim,F.A.,Houssein,E.H.,Mabrouk,M.S.,Al-Atabany,W.& Mirjalili,S. (2024) Enhancing Parkinson’s disease diagnosis through stacked generalization of optimized machine learning models.Biomedical Signal Processing and Control,75,103590.
Hu,X.(2001) Using rough sets theory and database operations to construct a good ensemble of classifiers for data mining applications.Proceedings of the 2001 IEEE International Conference on Data Mining,233-240,IEEE.
Jafarzadeh,H.,Mahdianpari,M.,Gill,E.,Mohammadimanesh,F.& Homayouni,S. (2021) Bagging and boosting ensemble classifiers for classification of multispectral, hyperspectral and PolSAR data: A comparative evaluation.Remote Sensing,13(21),4405.
Johnson, P., Vandewater, L., Wilson, W., Maruff, P., Savage, G., Graham, P., Macaulay, L. S., Ellis, K. A., Szoeke, C., & Martins, R. N. (2014). Genetic algorithm with logistic regression for prediction of progression to Alzheimer's disease. BMC Bioinformatics, 15, 1-14.
Ke,G.,Meng,Q.,Finley,T.,Wang,T.,Chen,W.,Ma,W.,Ye,Q.& Liu,T.-Y. (2017) LightGBM: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems,31,3149–3157.
Kennedy, J. & Eberhart, R. (1995). Particle swarm optimization. Proceedings of ICNN'95-International Conference on Neural Networks, 4, 1942-1948.
Liu,C.& Ji,L.(2025) Transformer fault diagnosis based on parallel AdaBoost-NB algorithm on Spark cloud platform. International Journal of Grid and High Performance Computing,17(1),1-23.
Mienye,I.D.& Sun,Y.(2022) A survey of ensemble learning: Concepts, algorithms, applications, and prospects.IEEE Access,10,122405–122428.
Miguel-Hurtado,O.,Guest,R.,Stevenage,S.V.,Neil,G.J.& Black,S.(2016) Comparing machine learning classifiers and linear/logistic regression to explore the relationship between hand dimensions and demographic characteristics.PLoS ONE,11(11),e0165521.
Sagi,O.& Rokach,L.(2018) Ensemble learning: A survey.WIREs Data Mining and Knowledge Discovery,8(4),e1249.
Schapire,R.E.(1990) The strength of weak learnability.Machine Learning,5,197–227.
Sevinç,E.(2022) An empowered AdaBoost algorithm implementation: A COVID-19 dataset study.Computers & Industrial Engineering,165, 107912.
Tan,P.N.,Steinbach,M.,Karpatne,A.& Kumar,V.(2019) Introduction to data mining (2nd Edition).Pearson Education.
Velarde,G.,Sudhir,A.,Deshmane,S.,Deshmunkh,A.,Sharma,K.& Joshi,V. (2023) Evaluating XGBoost for balanced and imbalanced data: Application to fraud detection. Paper presented at NVIDIA GTC 2023–The Conference for the Era of AI and the Metaverse.
Wang,R.(2012) AdaBoost for feature selection, classification and its relation with SVM: A review.Physics Procedia,25,800–807.
Wiens,M.,Verone-Boyle,A.Henscheid,N.,Podichetty,J.T.& Burton,J. (2025) A tutorial and use case example of the eXtreme gradient boosting (XGBoost) artificial intelligence algorithm for drug development applications. Clinical and Translational Science,18,e70172.
Wolpert,D.H.(1992) Stacked generalization.Neural Networks,5(2),241–259.
Zhou,T.& Jiao,H.(2023) Exploration of the stacking ensemble machine learning algorithm for cheating detection in large-scale assessment.Educational and Psychological Measurement,83(4),831–854.
Zizaan,A.& Idri,A.(2024) Evaluating and comparing bagging and boosting of hybrid learning for breast cancer screening. Scientific African,23,e01989.
Zuo,D.,Yang,L.,Jin,Y.,Qi,H.,Liu,Y.& Ren,L.(2023) Machine learning-based models for the prediction of breast cancer recurrence risk. BMC Medical Informatics and Decision Making,23,276.