| 研究生: |
吳善旋 Wu, Shan-Hsuan |
|---|---|
| 論文名稱: |
應用類神經網路及決策樹於肝癌罹患預測分析之研究 A Study on Predicting the Liver Cancer by the Artificial Neural Network and the Decision Tree |
| 指導教授: |
陳澤生
Chen, Tse-Sheng |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 工程管理碩士在職專班 Engineering Management Graduate Program |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 中文 |
| 論文頁數: | 76 |
| 中文關鍵詞: | 類神經網路 、決策樹 、肝癌 |
| 外文關鍵詞: | Artificial Neural Network, Decision Tree, Liver Cancer |
| 相關次數: | 點閱:197 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
依據行政院衛生署的資料,在台灣,二十多年來,肝癌一直位於全國十大癌症死因前三名,因肝癌而死亡的人數每年達約5,000至7,000人。因應資料探勘技術興起,利用大數據建立之預測模型可提供實證醫學於臨床研究之應用,本研究使用臨床上常見的生化檢測數據做為特徵變數,建構預測模型,篩檢出潛在肝癌病患、提供醫療決策評估輔助。
本研究之資料來自Kaggle資料庫,原始資料共有583筆,使用R語言(R Studio軟體) 進行資料預處理、填補遺漏值後,583筆資料均可作為資料集使用,其中有416筆病患及167筆非病患,資料集共有10個特徵屬性,均可作為臨床上判斷肝臟狀況的特徵指數。本研究使用類神經網路及決策樹兩個探勘方法建立預測模型,將資料集分為80%訓練集與20%測試集進行實作和分析。經過機器學習及參數調整後,類神經網路及決策樹預測模型之準確度分別為77.8%及78.6%。
綜觀本研究所建立的模型及預測結果,說明可以使用有較少的檢查數據及特徵項目,運用資料探勘技術建立預測模型,針對特定的重大疾病做出預測,作為醫療決策之輔助使用。
According to data from the Executive Yuan Department of Health, the liver cancer has been taken up a position among the top three causes of cancer deaths in the country for more than 20 years. The number of deaths is about 5,000 to 7,000 per year. In response to the rise of data exploration technology, a predictive model has been built by using big data which can provide empirical medicine for clinical research applications. This study uses clinically common biochemical test data as characteristic variables for constructing predictive models to screen and forcearte the potential liver cancer patients, provide medical decision-making evaluation assistance.
The data source of this study is supported by the Kaggle database. There are a total of 583 original data. After using R language (R Studio software) for data preprocessing and filling in missing values, 583 data can be used as a data set with 416 patients and 167 non-patients. The data set has 10 characteristic attributes, which can be used as a characteristic index for clinical judgment of liver cancer condition. In this study, two exploration methods, artificial neural network and decision tree, were used to establish a prediction model, and the data set was divided into 80% training set and 20% test set for implementation and analysis. After machine learning and parameter adjustment, the accuracy of the artificial neural network and decision tree prediction models are 77.8% and 78.6%, respectively.
A comprehensive review of the model and prediction results established in this research shows that it is possible to use fewer inspection data and feature items and use data exploration technology to establish a prediction model to make predictions for specific major diseases as a supplement for medical decision-making.
中文部份
1. 江素嬌 (2010)。B 型肝炎帶原者接受追蹤檢查行為影響因素探討-以某基金會肝炎篩檢活動帶原者為例。國立臺灣師範大學健康促進與衛生教育學系研究所碩士論文,臺北市,臺灣。
2. 吳杰成 (2014)。運用機器學習演算法對脂肪肝預測研究。國立臺北醫學大學醫學資訊研究所碩士論文,臺北市,臺灣。
3. 吳瓊芳 (2014)。利用類神經網路預測肝癌患者射頻燒灼術後無病生存期。國立臺北醫學大學醫學資訊研究所碩士論文,臺北市,臺灣。
4. 吳耀銘與李伯皇(2006)。肝癌。台灣醫學,第10卷,第4期,pp. 482-487,臺北市,臺灣。
5. 李仁鐘 (2015)。應用 R 語言於資料分析-從機器學習、資料探勘到巨量資料 (第一版) 。臺北市,臺灣:松崗圖書。
6. 李俊宏與古清仁 (2010)。類神經網路與資料探勘技術在醫療診斷之應用研究。工程科技與教育學刊,第7卷,第1期,pp. 154-169,高雄市,臺灣
7. 林志陵與高嘉宏(2008)。肝癌的流行病學。中華民國癌症醫學會雜誌,第24卷,第5期,pp. 277-281,臺北市,臺灣。
8. 林宥安 (2014)。以健保資料庫與癌症登記檔建構糖尿病確診後罹患為肝癌之預測模型。國立臺北醫學大學醫學資訊研究所碩士論文,臺北市,臺灣。
9. 林錦成 (2010)。目前台灣地區肝癌流行與醫療趨勢。北市中醫會刊,第16卷,第3期,pp. 75-96,臺北市,臺灣。
10. 段名聰 (2020) 。運用資料探勘技術由健康檢查與生活習慣資料建立冠狀動脈心臟病預測模型。國立成功大學工程管理碩士在職專班碩士論文,臺南市,臺灣。
11. 馬芳資與林我聰 (2003)。決策樹形式知識之線上預測系統架構。圖書館學與資訊科學,第29卷,第2期,pp. 60-76,臺北市,臺灣。
12. 梁家維 (2018)。利用深度學習以診斷及用藥歷史預測罹癌風險-以肝癌為例。國立臺北醫學大學醫學資訊研究所碩士論文,臺北市,臺灣。
13. 許偉帆 (2017)。肝臟損傷指標 GOT與GPT的判讀。中國醫訊,166,pp. 11-12,臺中市,臺灣。
14. 陳正美、徐建業、邱泓文、白其卉與吳柏動 (2011)。以類神經網路及分類迴歸樹輔助肝癌病患預測存活情形。台灣公共衛生雜誌,第30卷,第5期,pp. 481-493,臺北市,臺灣。
15. 陳志華、楊子緯、張訓楨與賴永崧(2016)。特徵分析和機器學習方法應用於肝臟疾病檢測。福祉科技與服務管理學刊,第4卷,第3期,pp. 417-429,桃園市,臺灣。
16. 楊培銘 (2015)。沉默殺手肝癌。愛Care,65,pp. 6-7,臺北市,台灣。
17. 廖述賢 (2007)。資訊管理(第一版)。臺北市,臺灣:雙葉書廊。
18. 廖述賢與溫志皓 (2019)。資料探勘:人工智慧與機器學習發展以SPSS Modeler為範例 (第一版)。新北市,臺灣:博碩。
19. 劉振隆與陳建良 (2016)。以決策樹分析與模糊邏輯技術建立肝病罹病風險評估系統。Electronic Commerce Studies,第4卷,第3期,pp. 475-498,新北市,台灣。
20. 劉嘉玲、張秀芳、黃志傑與周玉民 (2016)。臺灣 B 型肝炎防治。疫情報導,第32卷,第14期,pp.290-300,臺北市,臺灣。
21. 劉鐘軒、蔡正中與陳海雄 (2013)。肝癌的診斷及治療最新發展。內科學誌,第24卷,第2期,pp. 85-94,臺北市,臺灣。
22. 蔡蕙如、柯明中、張偉斌與劉德明 (2007)。應用類神經網路與分類迴歸樹於肝癌分類模式。北市醫學雜誌,第4卷,第8期,pp. 658-667,臺北市,臺灣。
23. 簡禎富與許嘉裕 (2014)。資料挖礦與大數據分析 (第一版)。臺北市,臺灣:前程文化。
英文部份
1. Agami, N., Atiya, A., Saleh, M., & El-Shishiny, H. (2009). A neural network based dynamic forecasting model for Trend Impact Analysis. Technological Forecasting and Social Change, 76(7), pp. 952-962, Nederland.
2. Bassi, P., Sacco, E., De Marco, V., Aragona, M., & Volpe, A. (2007). Prognostic accuracy of an artificial neural network in patients undergoing radical cystectomy for bladder cancer: a comparison with logistic regression analysis. BJU international, 99(5), pp. 1007-1012, USA.
3. Berry, M. J., & Linoff, G. S. (2004). Data mining techniques: for marketing, sales, and customer relationship management(2nd ed.).USA: John Wiley & Sons.
4. Bosch, F. X., Ribes, J., Díaz, M., & Cléries, R. (2004). Primary liver cancer: worldwide incidence and trends. Gastroenterology, 127(5), pp. 5-16, USA.
5. Breiman, L., Friedman, J., Stone, C., & Olshen, R. (1984). Classification and regression tree(1st ed.). United States: Chapman & Hall.
6. Chen, M.-S., Han, J., & Yu, P. S. (1996). Data mining: an overview from a database perspective. IEEE Transactions on Knowledge and data Engineering, 8(6), pp. 866-883, USA.
7. Davies, P. (1994). Design issues in neural network development. Neurovest Journal, 5, pp. 21-25, USA.
8. DeGroff, C. G., Bhatikar, S., Hertzberg, J., Shandas, R., Valdes-Cruz, L., & Mahajan, R. L. (2001). Artificial neural network–based method of screening heart murmurs in children. Circulation, 103(22), pp. 2711-2716, USA.
9. Delen, D., Walker, G., & Kadam, A. (2005). Predicting breast cancer survivability: a comparison of three data mining methods. Artificial intelligence in medicine, 34(2), pp.113-127,Netherlands.
10. Eslam, M., Hashem, A. M., Romero-Gomez, M., Berg, T., Dore, G. J., Mangia, A., . . . Abate, M. L. (2016). FibroGENE: a gene-based model for staging liver fibrosis. Journal of hepatology, 64(2), pp. 390-398, Switzerland.
11. Freeman, J.A., & Skapura, D. M. (1992). Neural Networks Algorithms, Applications, and Programming Techniques (1st ed). USA: Addison-Wesley Publishing Company.
12. Hsu, H.-M., Lu, C.-F., Lee, S.-C., Lin, S.-R., & Chen, D.-S. (1999). Seroepidemiologic survey for hepatitis B virus infection in Taiwan: the effect of hepatitis B mass immunization. The Journal of infectious diseases, 179(2), pp. 367-370, USA.
13. Fayyad, U., Piatetsky-Shapiro, G., & Smyth, P. (1996). From data mining to knowledge discovery in databases. AI magazine, 17(3), pp. 37-37, USA.
14. Hall, C. (1995). The devil's in the details: techniques, tools, and application for database mining and knowledge discovery part II. Intelligent Software Strategies, 6(9), pp. 1-16, USA.
15. Hsu, H.-M., Lu, C.-F., Lee, S.-C., Lin, S.-R., & Chen, D.-S. (1999). Seroepidemiologic survey for hepatitis B virus infection in Taiwan: the effect of hepatitis B mass immunization. The Journal of infectious diseases, 179(2), pp. 367-370, UK.
16. Konerman, M. A., Beste, L. A., Van, T., Liu, B., Zhang, X., Zhu, J., . . . Ioannou, G. N. (2019). Machine learning models to predict disease progression among veterans with hepatitis C virus. Plos one, 14(1), USA.
17. Law, R. (2000). Back-propagation learning in improving the accuracy of neural network-based tourism demand forecasting. Tourism Management, 21(4), pp. 331-340, Nederland.
18. Sahoo, G. B., Ray, C., Wang, J. Z., Hubbs, S. A., Song, R., Jasperse, J., & Seymour, D. (2005). Use of artificial neural networks to evaluate the effectiveness of riverbank filtration. Water Research, 39(12), pp. 2505-2516, USA.
19. Siegel, R. L., Miller, K. D., & Jemal, A. (2019). Cancer statistics, 2019. CA: a cancer journal for clinicians, 69(1), pp. 7-34, USA.
20. Seliya, N., Khoshgoftaar, T. M., & Van Hulse, J. (2009). A study on the relationships of classifier performance metrics. 21st IEEE International Conference on Tools with Artificial Intelligence, New Jersey, USA.
網站部份
1. 台灣癌症防治網站(2012)。許晉輝:如何看肝功能的檢查報告。取自:http://web.tccf.org.tw/lib/addon.php?act=post&id=3022。最後瀏覽日:2020/11/21
2. 衛生福利部網站 (2020)。統計處:108年國人死因統計結果。取自:https://www.mohw.gov.tw/cp-16-54482-1.html。最後瀏覽日:2020/11/14
3. 泛科學網站(2019)。吳育瑋:人工智慧的「黑箱作業」,類神經網路如何將生物分類的。取自:https://pansci.asia/archives/163315。最後瀏覽日:2021/03/19
4. Towards data science (2017)。Prashant Gupta:Decision Trees in Machine Learning。取自:https://towardsdatascience.com/decision-trees-in-machine-learning-641b9c4e8052。最後瀏覽日:2021/04/17
5. Kaggle (2017)。JeevanNagaraj:Indian Liver Patient Dataset。取自:https://www.kaggle.com/jeevannagaraj/indian-liver-patient-dataset。最後瀏覽日:2021/04/19
6. 衛生福利部中央健康保險署網站(2019)。健保署:掌握肝癌風險‼健康存摺「肝癌風險預測」來幫你。取自:https://www.nhi.gov.tw/Nhi_E-LibraryPubWeb/Epaper/ItemDetail.aspx?DataID=5390&IsWebData=0&ItemTypeID=5&PapersID=515&PicID=。最後瀏覽日:2021/05/22
7. IEEE Spectrum (2018)。Stephen Cass:The 2018 Top Programming Languages。取自:https://spectrum.ieee.org/at-work/innovation/the-2018-top-programming-languages。最後瀏覽日:2021/05/22