簡易檢索 / 詳目顯示

研究生: 許瑞允
Hsu, Rui-Yun
論文名稱: 機器學習於台積電月營收年增率預測之應用:多模型比較分析
An Application of Machine Learning to Forecasting TSMC's Monthly Revenue Growth: A Multi-Model Comparative Analysis
指導教授: 顏盟峯
Yen, Meng-Feng
學位類別: 碩士
Master
系所名稱: 管理學院 - 財務金融研究所
Graduate Institute of Finance
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 119
中文關鍵詞: 台積電營收預測機器學習XGBoost隨機森林
外文關鍵詞: TSMC, revenue forecasting, machine learning, XGBoost, Random Forest
ORCID: ORCID iD: 0000-0002-6989-2786
相關次數: 點閱:4下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 人工智慧與高效能運算需求快速擴張,使半導體產業在全球科技供應鏈中的重要性持續提升。台積電作為全球晶圓代工龍頭,其月營收不僅反映公司營運成果,也常被視為觀察半導體景氣循環與科技投資動能的重要訊號。因此,建立兼具準確度、穩健性與可解釋性之營收預測模型,對投資決策、企業評價及產業分析均具有實務意義。
    本研究以台積電為研究對象,以月營收年增率為預測目標,樣本涵蓋2014至2025年,並以人工智慧需求快速成長之2024至2025年作為樣本外測試期。解釋變數涵蓋公司內部財務指標、半導體產業景氣指標、美元兌新台幣匯率、半導體市場代理變數,以及外銷訂單等總體經濟領先指標;模型方面則比較線性迴歸、隨機森林、XGBoost、LSTM與SARIMAX五種方法之樣本外預測表現,研究設計並嚴格區分訓練與測試資訊,以降低資料洩漏與前視偏誤。
    實證結果顯示,樹狀整合模型整體優於線性迴歸、LSTM與SARIMAX。納入產業景氣與市場代理變數後,XGBoost於固定訓練/測試切分下表現最佳;於逐月更新之滾動驗證下,隨機森林則展現較佳之跨期穩健性,兩者差異未達統計顯著,且均明顯優於簡單之天真法基準。進一步加入其他總體經濟指標並未帶來額外改善,顯示具代表性且及時之市場訊號,較變數數量之多寡更為重要。解釋性分析亦指出,資本支出與研發投入之遞延效果、產業景氣、匯率與毛利率,均為重要之預測訊號。
    惟2024至2025年之營收高成長型態與訓練期資料分布明顯不同,使各模型均面臨外推限制,樹狀模型於部分高成長月份仍可能低估。整體而言,XGBoost與隨機森林互有所長,可共同作為台積電月營收即時推估(nowcasting)之主要工具。

    The rapid expansion of artificial intelligence (AI) and high-performance computing has steadily raised the importance of the semiconductor industry in the global technology supply chain. As the world's leading foundry, Taiwan Semiconductor Manufacturing Company (TSMC) reports monthly revenue that reflects not only its own operating performance but also broader semiconductor cycles and the momentum in technology investment. A revenue forecasting model that combines accuracy, robustness, and interpretability is therefore of practical value for investment decisions, corporate valuation, and industry analysis.

    This study uses TSMC as its subject and forecasts its year-over-year revenue growth. The sample covers 2014 to 2025, with the AI-driven high-growth period of 2024-2025 reserved as the out-of-sample test period. The predictors include internal financial indicators, semiconductor industry-cycle indicators, the USD/TWD exchange rate, a semiconductor market proxy, and macroeconomic leading indicators such as export orders. Five models are compared: linear regression, Random Forest, XGBoost, LSTM, and SARIMAX. Training and test information are strictly separated to reduce data leakage and look-ahead bias.

    The results show that tree-based ensemble models outperform linear regression, LSTM, and SARIMAX overall. Once industry-cycle and market-proxy variables are included, XGBoost performs best under the fixed split, while Random Forest exhibits greater cross-period robustness under month-by-month rolling validation; the difference between the two is not statistically significant, and both clearly outperform naive benchmarks. Adding further macroeconomic indicators brings no additional improvement, suggesting that representative and timely market signals matter more than the number of variables. Interpretability analysis identifies the delayed effects of capital expenditure and R&D investment, industry-cycle conditions, the exchange rate, and gross margin as important predictive signals.

    Because revenue growth in 2024-2025 differs markedly from the training-period distribution, all models face extrapolation limits, and tree-based models may still underestimate some high-growth months. Overall, XGBoost and Random Forest have complementary strengths and together serve as practical tools for nowcasting TSMC's monthly revenue.

    摘 要 I Extended Abstract II 誌謝 VII 目錄 VIII 圖目錄 XI 表目錄 XII 第一章 緒論 1 1.1 研究背景 1 1.2 研究動機 2 1.3 研究目的 2 1.4 研究貢獻 4 第二章 文獻探討 5 2.1 營收預測之理論基礎與研究意義 5 2.1.1 營收在企業財務分析中的重要性 5 2.1.2 營收預測對投資決策與企業評價之意義 6 2.1.3 傳統營收預測方法之發展 7 2.2 傳統預測模型於營收預測之應用 8 2.2.1 時間序列模型:以 SARIMAX 為例 8 2.2.2 多元線性迴歸模型 10 2.2.3 傳統模型之限制 12 2.3 機器學習於財務預測之應用 13 2.3.1 機器學習預測之基本概念 13 2.3.2 線性迴歸在財務預測中的角色 14 2.3.3 隨機森林(Random Forest)於預測問題之應用 14 2.3.4 極限梯度提升樹(XGBoost)之應用 16 2.3.5 長短期記憶網路(LSTM)之應用 18 2.3.6 機器學習模型相較於傳統模型之優勢與限制 20 2.4 影響企業營收之因素 21 2.4.1 公司內部財務指標對營收之影響 21 2.4.2 外部市場因素對營收之影響 22 2.5 半導體產業特性與台積電營收預測之特殊性 25 2.6 文獻回顧小結與研究缺口 25 第三章 研究方法 28 3.1 研究架構與流程 28 3.2 研究對象與資料來源 29 3.3 研究變數定義 30 3.3.1 目標變數定義 30 3.3.2 解釋變數分類 31 3.3.3 落後變數設計 33 3.4 資料前處理與資料特性診斷 34 3.4.1 變數敘述統計 35 3.4.2 相關係數與多重共線性診斷 37 3.4.3 平穩性檢定 39 3.5 模型建構 41 3.5.1 線性迴歸模型(Linear Regression) 42 3.5.2 隨機森林模型(Random Forest) 43 3.5.3 極限梯度提升樹模型(XGBoost) 44 3.5.4 長短期記憶網路(LSTM) 46 3.5.5 季節性自迴歸整合移動平均模型(SARIMAX) 47 3.6 模型訓練與預測流程 48 3.7 模型評估指標 50 3.7.1 平均絕對百分比誤差(Mean Absolute Percentage Error, MAPE) 50 3.7.2 調整後判定係數(Adjusted R²) 51 3.7.3 均方根誤差(RMSE) 51 3.8 模型設定參數 51 第四章 結果與討論 56 4.1 情境一:僅內部財務指標之預測結果 56 4.2 情境二:再納入 SOX 產業景氣變數後之預測結果 58 4.3 情境三:進一步納入半導體市場代理變數後之預測結果 60 4.4 情境四:加入總體經濟領先指標之影響 62 4.5 各情境最佳模型之綜合比較 64 4.6 各情境下 Random Forest 與 XGBoost 之特徵重要性比較 67 4.7 各情境下 Random Forest 與 XGBoost 之 SHAP 分析 74 4.8 模型預測差異之顯著性檢定 82 4.9 預測目標形式之敏感度分析 83 4.10 逐步擴張視窗(recursive window)驗證 84 4.11 與天真基準(naïve benchmark)之比較 85 4.12 AI 結構性轉折下之模型穩健性探討 87 第五章 結論 89 5.1 研究結論 89 5.2 研究貢獻 91 5.3 研究限制 91 5.4 未來展望 92 參考文獻 93 附錄A 各情境最佳超參數彙整 95 附錄B 資料特性之補充診斷 96 附錄C 結構性轉折檢定之完整結果 99 附錄D 殘差自相關與預測區間診斷 101 附錄E 各情境完整特徵重要性數值 103 附錄F 程式碼與資料檔取得說明 105

    1. Aubry, M. and P. Renou-Maissant, Semiconductor industry cycles: Explanatory factors and forecasting. Economic Modelling, 2014. 39: p. 221-231.
    2. Mullainathan, S. and J. Spiess, Machine learning: an applied econometric approach. Journal of Economic Perspectives, 2017. 31(2): p. 87-106.
    3. Gu, S., B. Kelly, and D. Xiu, Empirical asset pricing via machine learning. The Review of Financial Studies, 2020. 33(5): p. 2223-2273.
    4. Ghosh, A., Z. Gu, and P.C. Jain, Sustained earnings and revenue growth, earnings quality, and earnings response coefficients. Review of Accounting Studies, 2005. 10(1): p. 33-57.
    5. Penman, S.H., Financial statement analysis and security valuation. 2010, New York: McGraw-Hill/Irwin.
    6. Ohlson, J.A., Earnings, book values, and dividends in equity valuation. Contemporary Accounting Research, 1995. 11(2): p. 661-687.
    7. Nissim, D. and S.H. Penman, Ratio analysis and equity valuation: From research to practice. Review of Accounting Studies, 2001. 6(1): p. 109-154.
    8. Box, G.E.P., G.M. Jenkins, G.C. Reinsel, and G.M. Ljung, Time series analysis: forecasting and control. 2015, Hoboken, NJ: John Wiley & Sons.
    9. Bradshaw, M.T., Analysts’ forecasts: what do we know after decades of work? 2011, SSRN Working Paper No. 1880339.
    10. Hyndman, R.J. and G. Athanasopoulos, Forecasting: principles and practice. 2018, Melbourne, Australia: OTexts.
    11. Wooldridge, J.M., Introductory econometrics: a modern approach. 2016, Boston, MA: South-Western Cengage Learning.
    12. Breiman, L., Random forests. Machine Learning, 2001. 45(1): p. 5-32.
    13. Schonlau, M. and R.Y. Zou, The random forest algorithm for statistical learning. The Stata Journal, 2020. 20(1): p. 3-29.
    14. Friedman, J.H., Greedy function approximation: a gradient boosting machine. Annals of Statistics, 2001. 29(5): p. 1189-1232.
    15. Hastie, T., R. Tibshirani, and J. Friedman, The elements of statistical learning: data mining, inference, and prediction. 2009, New York: Springer.
    16. Chen, T. and C. Guestrin, XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016: p. 785-794.
    17. Hochreiter, S. and J. Schmidhuber, Long short-term memory. Neural Computation, 1997. 9(8): p. 1735-1780.
    18. McConnell, J.J. and C.J. Muscarella, Corporate capital expenditure decisions and the market value of the firm. Journal of Financial Economics, 1985. 14(3): p. 399-422.
    19. Park, C.W. and M. Pincus, A reexamination of the incremental information content of capital expenditures. Journal of Accounting, Auditing & Finance, 2003. 18(2): p. 281-302.
    20. Lev, B. and T. Sougiannis, The capitalization, amortization, and value-relevance of R&D. Journal of Accounting and Economics, 1996. 21(1): p. 107-138.
    21. García-Manjón, J.V. and M.E. Romero-Merino, Research, development, and firm growth. Empirical evidence from European top R&D spending firms. Research Policy, 2012. 41(6): p. 1084-1092.
    22. Dominguez, K.M. and L.L. Tesar, Exchange rate exposure. Journal of International Economics, 2006. 68(1): p. 188-218.
    23. Liu, W.-H., Determinants of the semiconductor industry cycles. Journal of Policy Modeling, 2005. 27(7): p. 853-866.
    24. Babina, T., A. Fedyk, A. He, and J. Hodson, Artificial intelligence, firm growth, and product innovation. Journal of Financial Economics, 2024. 151: p. 103745.
    25. Steinmeister, L. and M. Pauly, Human vs. machines: Who wins in semiconductor market forecasting? Expert Systems with Applications, 2025. 263: p. 125719.
    26. Lundberg, S.M. and S.-I. Lee, A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 2017. 30.
    27. Diebold, F.X. and R.S. Mariano, Comparing predictive accuracy. Journal of Business & Economic Statistics, 1995. 13(3): p. 253-263.
    28. Harvey, D., S. Leybourne, and P. Newbold, Testing the equality of prediction mean squared errors. International Journal of Forecasting, 1997. 13(2): p. 281-291.

    QR CODE