簡易檢索 / 詳目顯示

研究生: 董建維
Dong, Jian-Wei
論文名稱: 基於多尺度填息機率估計與強化學習門控機制之除權息事件交易策略
An Ex-Dividend Event Trading Strategy Based on Multi-Scale Fill Probability Estimation and a Reinforcement Learning-Based Meta-Gating Mechanism
指導教授: 陳朝鈞
Chen, Chao-Chun
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 127
中文關鍵詞: 除息事件選擇性執行回溯性診斷覆蓋率分解可靠度資訊閘門決策支援
外文關鍵詞: Ex-dividend events, Selective execution, Retrospective diagnostics, Coverage decomposition, Reliability-informed gating, Decision support
相關次數: 點閱:4下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本論文探討金融人工智慧系統中由預測分數轉換為實際執行決策的中介層,並以臺灣上市股票除息事件作為研究情境。傳統金融機器學習流程常以固定門檻將模型分數直接轉為交易訊號,但在非定態市場中,模型的排序能力與分數可靠度可能隨時間改變,因此高分並不必然代表訊號在當下仍值得執行。為處理此問題,本研究提出模組化的「分數–篩選–閘門」回溯性診斷架構,將預測與執行控制分離,並把接受覆蓋率明確納入評估。

    第一階段的填息分數估計器(Recovery Score Estimator, RSE)使用除息前 20 個交易日的價格序列產生有界填息分數,再以固定門檻篩選候選事件。輸入表示比較原始價格與完整集合經驗模態分解搭配自適應雜訊(CEEMDAN)所形成的多尺度分支,預測骨幹則包含 LSTM、GRU、Transformer 與 Random Forest。第二階段的可靠度資訊閘門(Reliability-Informed Gate, RIG)以近端策略最佳化(Proximal Policy Optimization, PPO)訓練,並根據目前分數、與篩選邊界的距離、回溯更新的預測正確率與 Brier 誤差摘要、填息率、股利殖利率、已處理紀錄數及事件報酬變異等資訊,決定接受或抑制候選事件。

    實驗資料涵蓋 2015 至 2024 年,主要實驗採用依時間排序的四階段設計:2015–2018 年訓練 RSE、2019–2020 年進行 RSE 驗證、2021–2022 年訓練 RIG、2023–2024 年作為最終測試,並比較六種持有期間。第一階段模型的 AUC 僅介於 0.5019–0.5318,顯示預測訊號偏弱。在合併固定名目金額的回溯性診斷中,RIG 接受 17,782 個第一階段候選中的 3,519 個,接受覆蓋率為 19.79%。相較於直接執行所有候選事件,總加總貢獻由 −0.3518 變為 0.0054,累積固定名目貢獻的峰谷下降由 −0.4113 變為 −0.0618。

    本研究的主要貢獻是把填息預測、候選篩選與交易曝險控制分開,並利用覆蓋率分解區分候選數量、接受覆蓋率、條件式接受事件報酬與總加總貢獻,避免因降低交易參與而產生的機械性改善被誤解為選擇能力提升。由於目前可靠度狀態使用最終結果進行回溯更新、不同實驗組態的事件紀錄會被合併、參考價格並非訊號形成後可直接成交的價格,且主要比較流程具有不同覆蓋率,因此本文結果應解讀為回溯性曝險控制診斷,而不是線上因果控制、PPO 優越性、可執行獲利或資本約束投資組合績效的證明。

    This thesis investigates the intermediate decision layer that converts a supervised financial-model score into an execution decision. The empirical setting is ex-dividend events in Taiwan-listed stocks. A modular score–screen–gate framework is developed in which a recovery score estimator first identifies candidates and a reliability-informed gate subsequently accepts or suppresses those candidates. The design makes acceptance coverage explicit so that the economic effect of executing fewer opportunities can be distinguished from changes in the composition of accepted events. The empirical evidence should be interpreted as a retrospective diagnostic rather than a causal online trading experiment because reliability summaries are updated from eventual outcomes and the reported event payoffs use a pre-event reference-price convention.

    中文摘要 I Abstract III 誌謝 X 目錄 XI 表目錄 XV 圖目錄 XVII 符號說明 XIX 第一章 緒論 1 1-1. 研究背景 1 1-2. 研究動機 3 1-3. 研究問題 5 1-4. 研究目的 7 1-5. 研究方法概述 8 1-6. 研究貢獻 9 1-7. 研究範圍與限制 10 1-8. 論文架構 11 第二章 文獻回顧 13 2-1. 除權息與填息行為 13 2-2. 事件研究與事件型交易 15 2-3. 機器學習於金融時間序列預測 16 2-4. 時間序列分解與多尺度金融訊號建模 18 2-5. 金融市場非平穩性與模型風險 20 2-6. 決策機制與強化學習於交易策略 21 2-7. 選擇性分類、棄權決策與覆蓋率觀點 23 2-8. 本章小結 24 第三章 研究方法 25 3-1. 問題定義(Problem Formulation) 25 3-1.1 成本調整後填息事件定義 26 3-1.2 填息分數估計與第一階段候選篩選 27 3-1.3 可靠度資訊閘門決策問題 36 3-1.4 交易建模目標、覆蓋率與選擇效果 37 3-2. 事件資料建構與資訊時間邊界 39 3-2.1 研究資料與事件視窗 39 3-2.2 決策日、結果可觀測日與同日事件 40 3-3. 方法架構與決策流程 41 3-4. 填息分數估計器(Recovery Score Estimator, RSE) 43 3-4.1 RSE 之角色與建模目標 43 3-4.2 CEEMDAN 方法概述 43 3-4.3 原始價格與多尺度輸入表示 44 3-4.4 Backbone 與編碼器架構 45 3-4.5 RSE 的限制與後續決策銜接 47 3-5. 可靠度資訊閘門(Reliability-Informed Gate, RIG) 48 3-5.1 RIG 之角色與狀態設計 48 3-5.2 動作、執行與獎勵設計 51 3-5.3 PPO 策略最佳化 52 3-5.4 RSE 與 RIG 之整合流程 53 3-6. 模型訓練、驗證與推論程序 53 3-6.1 第一階段 RSE 訓練與模型選擇 54 3-6.2 第二階段 RIG 訓練 55 3-6.3 最終推論 55 3-6.4 方法所估計的對象 55 第四章 實驗設計與實證分析 57 4-1. 實驗設定與參數 57 4-2. 基準模型、實驗變體與評估指標 59 4-3. 第一階段預測模型結果 61 4-4. 模型元件消融分析 62 4-4.1 Multiple Encoder 與 Single Encoder 63 4-4.2 CEEMDAN 與 RAW Price Sequence 65 4-5. 直接執行與 RIG 的選擇性執行結果 67 4-5.1 整體候選覆蓋率與合併結果 67 4-5.2 時間區間之結果異質性 68 4-5.3 跨模型與持有期間融合後的整體路徑 71 4-5.4 模型誤差時間聚集與候選抑制組成 74 4-5.5 不同持有期間之穩健性分析 76 4-5.6 產業別結果異質性 78 4-5.7 STOP 與 NO_STOP 的輔助比較 81 4-6. 第一階段分數可靠度與相同覆蓋率診斷 82 4-7. 本章小結 83 第五章 結論與未來研究方向 85 5-1. 研究結果總結 85 5-2. 研究問題回應 87 5-3. 研究貢獻與方法學意涵 90 5-4. 研究限制 91 5-5. 未來研究方向 92 5-6. 結論 93 參考文獻 94 附錄 A:未來相同覆蓋率研究之延伸設計 98 附錄 B:未來重現研究建議保存之中介資料 99 附錄 C:重現設定與超參數 100

    [1] Andrew Ang and Allan Timmermann. Regime changes and financial markets. Annual Review of Financial Economics, 4:313–337, 2012.

    [2] Shaojie Bai, J. Zico Kolter, and Vladlen Koltun. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271, 2018.

    [3] Tim Bollerslev. Generalized autoregressive conditional heteroskedasticity. Journal of Econometrics, 31:307–327, 1986.

    [4] George E. P. Box, Gwilym M. Jenkins, Gregory C. Reinsel, and Greta M. Ljung. Time Series Analysis: Forecasting and Control. John Wiley & Sons, Hoboken, NJ, 5 edition, 2015.

    [5] Leo Breiman. Random forests. Machine Learning, 45:5–32, 2001.

    [6] Stephen J. Brown and Jerold B. Warner. Using daily stock returns: The case of event studies. Journal of Financial Economics, 14:3–31, 1985.

    [7] Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder–decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, pages 1724–1734, Doha, Qatar, 2014.

    [8] Rama Cont. Model uncertainty and its impact on the pricing of derivative instruments. Mathematical Finance, 16(3):519–547, 2006.

    [9] Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine Learning, 20:273–297, 1995.

    [10] Thomas G. Dietterich. Ensemble methods in machine learning. In Multiple Classifier Systems, pages 1–15, Berlin, Heidelberg, 2000. Springer.

    [11] Ran El-Yaniv and Yair Wiener. On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11:1605–1641, 2010.

    [12] Edwin J. Elton and Martin J. Gruber. Marginal stockholder tax rates and the clientele effect. The Review of Economics and Statistics, 52:68–74, 1970.

    [13] Eugene F. Fama. Efficient capital markets: A review of theory and empirical work. The Journal of Finance, 25:383–417, 1970.

    [14] Jerome H. Friedman. Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29:1189–1232, 2001.

    [15] Yonatan Geifman and Ran El-Yaniv. Selective classification for deep neural networks. In Advances in Neural Information Processing Systems, pages 4878–4887, 2017.

    [16] Richard C. Green. Taxation and the ex-dividend day behavior of common stock prices. Journal of Financial Economics, 1980.

    [17] Sepp Hochreiter and Jurgen Schmidhuber. Long short-term memory. Neural Computation, 9:1735–1780, 1997.

    [18] A. C. Hsu and S. H. Lin. Trading strategies based on dividend yield: Evidence from the taiwan stock market. The International Journal of Business and Finance Research, 4:71–84, 2010.

    [19] Shing-Yang Hu and Yung-Il Tseng. Who wants to trade around ex-dividend days? Financial Management, 35:95–119, 2006.

    [20] Norden E. Huang, Zheng Shen, Steven R. Long, Manli C. Wu, Hsing H. Shih, Quanan Zheng, Nai-Chyuan Yen, Chi Chao Tung, and Henry H. Liu. The empirical mode decomposition and the hilbert spectrum for nonlinear and non-stationary time series analysis. Proceedings of the Royal Society of London. Series A, 454:903–995, 1998.

    [21] Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton. Adaptive mixtures of local experts. Neural Computation, 3:79–87, 1991.

    [22] Andrew W. Lo and A. Craig MacKinlay. Stock market prices do not follow random walks: Evidence from a simple specification test. Review of Financial Studies, 1(1):41–66, 1988.

    [23] A. Craig MacKinlay. Event studies in economics and finance. Journal of Economic Literature, 35:13–39, 1997.

    [24] Stephane G. Mallat. A theory for multiresolution signal decomposition: The wavelet representation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 11(7):674–693, 1989.

    [25] R. Tyrrell Rockafellar and Stanislav Uryasev. Optimization of conditional value-at-risk. The Journal of Risk, 2:21–41, 2000.

    [26] Suman Saha, Junbin Gao, and Richard Gerlach. Stock movement prediction on ex-dividend day using event-specific features and machine learning techniques. In International Joint Conference on Neural Networks, 2021.

    [27] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms, 2017.

    [28] Omer Berat Sezer, Mehmet Ugur Gudelek, and Ahmet Murat Ozbayoglu. Financial time series forecasting with deep learning: A systematic literature review: 2005–2019. Applied Soft Computing, 90:106181, 2020.

    [29] Shuo Sun, Rong Wang, and Bo An. Reinforcement learning for quantitative trading. ACM Transactions on Intelligent Systems and Technology, 14:1–29, 2023.

    [30] Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. The MIT Press, Cambridge, MA, 2 edition, 2018.

    [31] Maria E. Torres, Marcelo A. Colominas, Gaston Schlotthauer, and Patrick Flandrin. A complete ensemble empirical mode decomposition with adaptive noise. In 2011 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 4144–4147, Prague, Czech Republic, 2011.

    [32] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, pages 5998–6008, 2017.

    [33] Zhaohua Wu and Norden E. Huang. Ensemble empirical mode decomposition: A noise-assisted data analysis method. Advances in Adaptive Data Analysis, 1:1–41, 2009.

    QR CODE