| 研究生: |
許瑞允 Hsu, Rui-Yun |
|---|---|
| 論文名稱: |
機器學習於台積電月營收年增率預測之應用:多模型比較分析 An Application of Machine Learning to Forecasting TSMC's Monthly Revenue Growth: A Multi-Model Comparative Analysis |
| 指導教授: |
顏盟峯
Yen, Meng-Feng |
| 學位類別: |
碩士 Master |
| 系所名稱: |
管理學院 - 財務金融研究所 Graduate Institute of Finance |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 119 |
| 中文關鍵詞: | 台積電 、營收預測 、機器學習 、XGBoost 、隨機森林 |
| 外文關鍵詞: | TSMC, revenue forecasting, machine learning, XGBoost, Random Forest |
| ORCID: | ORCID iD: 0000-0002-6989-2786 |
| 相關次數: | 點閱:4 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
人工智慧與高效能運算需求快速擴張,使半導體產業在全球科技供應鏈中的重要性持續提升。台積電作為全球晶圓代工龍頭,其月營收不僅反映公司營運成果,也常被視為觀察半導體景氣循環與科技投資動能的重要訊號。因此,建立兼具準確度、穩健性與可解釋性之營收預測模型,對投資決策、企業評價及產業分析均具有實務意義。
本研究以台積電為研究對象,以月營收年增率為預測目標,樣本涵蓋2014至2025年,並以人工智慧需求快速成長之2024至2025年作為樣本外測試期。解釋變數涵蓋公司內部財務指標、半導體產業景氣指標、美元兌新台幣匯率、半導體市場代理變數,以及外銷訂單等總體經濟領先指標;模型方面則比較線性迴歸、隨機森林、XGBoost、LSTM與SARIMAX五種方法之樣本外預測表現,研究設計並嚴格區分訓練與測試資訊,以降低資料洩漏與前視偏誤。
實證結果顯示,樹狀整合模型整體優於線性迴歸、LSTM與SARIMAX。納入產業景氣與市場代理變數後,XGBoost於固定訓練/測試切分下表現最佳;於逐月更新之滾動驗證下,隨機森林則展現較佳之跨期穩健性,兩者差異未達統計顯著,且均明顯優於簡單之天真法基準。進一步加入其他總體經濟指標並未帶來額外改善,顯示具代表性且及時之市場訊號,較變數數量之多寡更為重要。解釋性分析亦指出,資本支出與研發投入之遞延效果、產業景氣、匯率與毛利率,均為重要之預測訊號。
惟2024至2025年之營收高成長型態與訓練期資料分布明顯不同,使各模型均面臨外推限制,樹狀模型於部分高成長月份仍可能低估。整體而言,XGBoost與隨機森林互有所長,可共同作為台積電月營收即時推估(nowcasting)之主要工具。
The rapid expansion of artificial intelligence (AI) and high-performance computing has steadily raised the importance of the semiconductor industry in the global technology supply chain. As the world's leading foundry, Taiwan Semiconductor Manufacturing Company (TSMC) reports monthly revenue that reflects not only its own operating performance but also broader semiconductor cycles and the momentum in technology investment. A revenue forecasting model that combines accuracy, robustness, and interpretability is therefore of practical value for investment decisions, corporate valuation, and industry analysis.
This study uses TSMC as its subject and forecasts its year-over-year revenue growth. The sample covers 2014 to 2025, with the AI-driven high-growth period of 2024-2025 reserved as the out-of-sample test period. The predictors include internal financial indicators, semiconductor industry-cycle indicators, the USD/TWD exchange rate, a semiconductor market proxy, and macroeconomic leading indicators such as export orders. Five models are compared: linear regression, Random Forest, XGBoost, LSTM, and SARIMAX. Training and test information are strictly separated to reduce data leakage and look-ahead bias.
The results show that tree-based ensemble models outperform linear regression, LSTM, and SARIMAX overall. Once industry-cycle and market-proxy variables are included, XGBoost performs best under the fixed split, while Random Forest exhibits greater cross-period robustness under month-by-month rolling validation; the difference between the two is not statistically significant, and both clearly outperform naive benchmarks. Adding further macroeconomic indicators brings no additional improvement, suggesting that representative and timely market signals matter more than the number of variables. Interpretability analysis identifies the delayed effects of capital expenditure and R&D investment, industry-cycle conditions, the exchange rate, and gross margin as important predictive signals.
Because revenue growth in 2024-2025 differs markedly from the training-period distribution, all models face extrapolation limits, and tree-based models may still underestimate some high-growth months. Overall, XGBoost and Random Forest have complementary strengths and together serve as practical tools for nowcasting TSMC's monthly revenue.
1. Aubry, M. and P. Renou-Maissant, Semiconductor industry cycles: Explanatory factors and forecasting. Economic Modelling, 2014. 39: p. 221-231.
2. Mullainathan, S. and J. Spiess, Machine learning: an applied econometric approach. Journal of Economic Perspectives, 2017. 31(2): p. 87-106.
3. Gu, S., B. Kelly, and D. Xiu, Empirical asset pricing via machine learning. The Review of Financial Studies, 2020. 33(5): p. 2223-2273.
4. Ghosh, A., Z. Gu, and P.C. Jain, Sustained earnings and revenue growth, earnings quality, and earnings response coefficients. Review of Accounting Studies, 2005. 10(1): p. 33-57.
5. Penman, S.H., Financial statement analysis and security valuation. 2010, New York: McGraw-Hill/Irwin.
6. Ohlson, J.A., Earnings, book values, and dividends in equity valuation. Contemporary Accounting Research, 1995. 11(2): p. 661-687.
7. Nissim, D. and S.H. Penman, Ratio analysis and equity valuation: From research to practice. Review of Accounting Studies, 2001. 6(1): p. 109-154.
8. Box, G.E.P., G.M. Jenkins, G.C. Reinsel, and G.M. Ljung, Time series analysis: forecasting and control. 2015, Hoboken, NJ: John Wiley & Sons.
9. Bradshaw, M.T., Analysts’ forecasts: what do we know after decades of work? 2011, SSRN Working Paper No. 1880339.
10. Hyndman, R.J. and G. Athanasopoulos, Forecasting: principles and practice. 2018, Melbourne, Australia: OTexts.
11. Wooldridge, J.M., Introductory econometrics: a modern approach. 2016, Boston, MA: South-Western Cengage Learning.
12. Breiman, L., Random forests. Machine Learning, 2001. 45(1): p. 5-32.
13. Schonlau, M. and R.Y. Zou, The random forest algorithm for statistical learning. The Stata Journal, 2020. 20(1): p. 3-29.
14. Friedman, J.H., Greedy function approximation: a gradient boosting machine. Annals of Statistics, 2001. 29(5): p. 1189-1232.
15. Hastie, T., R. Tibshirani, and J. Friedman, The elements of statistical learning: data mining, inference, and prediction. 2009, New York: Springer.
16. Chen, T. and C. Guestrin, XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016: p. 785-794.
17. Hochreiter, S. and J. Schmidhuber, Long short-term memory. Neural Computation, 1997. 9(8): p. 1735-1780.
18. McConnell, J.J. and C.J. Muscarella, Corporate capital expenditure decisions and the market value of the firm. Journal of Financial Economics, 1985. 14(3): p. 399-422.
19. Park, C.W. and M. Pincus, A reexamination of the incremental information content of capital expenditures. Journal of Accounting, Auditing & Finance, 2003. 18(2): p. 281-302.
20. Lev, B. and T. Sougiannis, The capitalization, amortization, and value-relevance of R&D. Journal of Accounting and Economics, 1996. 21(1): p. 107-138.
21. García-Manjón, J.V. and M.E. Romero-Merino, Research, development, and firm growth. Empirical evidence from European top R&D spending firms. Research Policy, 2012. 41(6): p. 1084-1092.
22. Dominguez, K.M. and L.L. Tesar, Exchange rate exposure. Journal of International Economics, 2006. 68(1): p. 188-218.
23. Liu, W.-H., Determinants of the semiconductor industry cycles. Journal of Policy Modeling, 2005. 27(7): p. 853-866.
24. Babina, T., A. Fedyk, A. He, and J. Hodson, Artificial intelligence, firm growth, and product innovation. Journal of Financial Economics, 2024. 151: p. 103745.
25. Steinmeister, L. and M. Pauly, Human vs. machines: Who wins in semiconductor market forecasting? Expert Systems with Applications, 2025. 263: p. 125719.
26. Lundberg, S.M. and S.-I. Lee, A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 2017. 30.
27. Diebold, F.X. and R.S. Mariano, Comparing predictive accuracy. Journal of Business & Economic Statistics, 1995. 13(3): p. 253-263.
28. Harvey, D., S. Leybourne, and P. Newbold, Testing the equality of prediction mean squared errors. International Journal of Forecasting, 1997. 13(2): p. 281-291.