簡易檢索 / 詳目顯示

研究生: 吳岳蓁
Wu, Yueh-Chen
論文名稱: 深度強化學習在序列稀疏獎勵任務中的系統性失敗診斷與 Multi-milestone PBRS設計原則
Systematic Failure Diagnosis and Multi-milestone PBRS Design Principles for Sequential Sparse Reward Tasks in Deep Reinforcement Learning
指導教授: 蕭宏章
Hsiao, Hung-Chang
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 59
中文關鍵詞: 深度強化學習 、稀疏獎勵 、潛力型獎勵塑形 、課程學習 、Double DQN 、條件必要性框架
外文關鍵詞: deep reinforcement learning, sparse reward, potential-based reward shaping, curriculum learning, Double DQN, failure diagnosis, Multi-milestone
相關次數: 點閱:143  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 序列稀疏獎勵任務要求 agent 依序完成多個子任務後方能獲得獎勵訊號,對深度強化學習構成根本性挑戰。本研究確立並實驗驗證了 Double DQN 在此類任務中失敗的主要瓶頸之一: TD 衰減瓶頸——約 100 步的任務鏈使最早子任務所接收之獎勵訊號衰減至 37% (𝛾^100 ≈ 0.37),為本實驗設定下之主導性失敗因素。
    針對此問題,本研究提出多里程碑潛力型獎勵塑形(Multi-milestone PBRS),於結構化任務檢查點植入進度訊號,並以單調性與安全性兩項設計約束加以規範。系統性消融分析確認 PBRS 為本實驗設定下之經驗性必要核心元件(無PBRS 配置:5/5 種子完全失敗),並於 DoorKey-16×16 達到穩定收斂(5/5 種子,mean_tail50 = 0.997)。
    跨任務驗證於 KeyCorridor-S3R3(5/5 種子收斂)表明,各設計元件之必要性取決於有效終端訊號強度(𝜉)而非固定設計,構成條件性必要框架。規模延伸實驗(DoorKey-8×8 及 KeyCorridor-S6R3,均 5/5 種子)進一步提供跨規模泛化之實證支持。

    Sequential sparse reward tasks—requiring ordered subtask completion before any reward signal becomes available—pose fundamental challenges for deep reinforcement learning. We identify and empirically validate a dominant failure mechanism in sequential sparse reward settings: a TD decay bottleneck, where an approximately 100-step task chain attenuates reward signals to 37% at the earliest subtask (𝛾^100 ≈ 0.37). We propose Multi-milestone Potential-Based Reward Shaping (Multi-milestone PBRS), which inserts immediate progress signals at structured task checkpoints under two formal design constraints—a monotonicity constraint and a safety constraint. Systematic ablation across 12 experimental configurations and 54 seed runs across four environments suggests that Double DQN + Multi-milestone PBRS constitutes a minimal empirically sufficient set for stable convergence (5/5 seeds under our experimental setting). Cross-task validation on KeyCorridor-S3R3 (5/5 seeds) reveals that component necessity is governed by effective terminal signal strength (𝜉) rather than a fixed prescription, forming a conditional necessity framework. Scale-extension experiments on DoorKey-8×8 and KeyCorridor-S6R3 (5/5 seeds each) further provide empirical support for cross-scale applicability.

    中文摘要 I Abstract II 誌謝 III Table of Contents IV List of Tables VI List of Figures VII List of Symbols VIII Chapter 1 Introduction 1 1-1 Research Background and Motivation 1 1-2 Limitations of Existing Approaches 1 1-3 Proposed Approach 2 1-4 Contributions 3 1-5 Thesis Organization 4 Chapter 2 Background and Related Work 6 2-1 Problem Formulation 6 2-2 Double DQN 6 2-3 Potential-Based Reward Shaping 7 2-4 Evaluation Environments 7 2-5 Related Work 10 2-5.1 PBRS-Related Work 10 2-5.2 Sparse Reward Methods 10 2-5.3 Algorithm Selection 12 2-5.4 Comparative Summary 12 Chapter 3 Failure Analysis 14 3-1 Root Cause: TD Decay Bottleneck 14 3-2 Terminal Signal Strength Analysis 15 3-2.1 Cross-Environment Comparison 15 3-2.2 Phase 2 as a Remediation Mechanism 16 3-2.3 Implications for KC-S6R3 17 Chapter 4 Multi-milestone PBRS: Method Design 18 4-1 PBRS Formulation 18 4-2 Milestone Potential Function Design 18 4-3 Two-Phase Curriculum Training 21 4-4 Optional Acceleration Component: goal-dynamic 22 Chapter 5 Experiments and Discussion 24 5-1 Experimental Setup 24 5-2 DoorKey-16×16: Main Method Results 25 5-3 Ablation Study on DoorKey-16×16 27 5-4 Phi Sensitivity Analysis 28 5-5 Comparison with Related Approaches 29 5-6 Cross-Task Generalization: KeyCorridor-S3R3 30 5-7 Scale Validation: DoorKey-8×8 32 5-8 Scale Extension: KeyCorridor-S6R3 34 5-9 Component Necessity Framework 36 Chapter 6 Conclusion and Future Work 40 6-1 Conclusion 40 6-2 Limitations 40 6-3 Method Applicability Framework 42 6-4 Future Work 43 References 45

    [1] Maxime Chevalier-Boisvert, Bolun Dai, Mark Towers, Rodrigo de Lazcano, Lucas Willems, Salem Lahlou, Suman Pal, Pablo Samuel Castro, and Jordan Terry. Minigrid & Miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks. In Advances in Neural Information Processing Systems: Datasets and Benchmarks Track, 2023.

    [2] Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. MIT Press, Cambridge, MA, 2nd edition, 2018.

    [3] Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba. Hindsight experience replay. In Advances in Neural Information Processing Systems (NeurIPS), volume 30, 2017.

    [4] Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. Curiosity-driven exploration by self-supervised prediction. In Proceedings of the International Conference on Machine Learning (ICML), pages 2778–2787, 2017.

    [5] Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov. Exploration by random network distillation. In Proceedings of the International Conference on Learning Representations (ICLR), 2019.

    [6] Sanmit Narvekar, Bei Peng, Matteo Leonetti, Jivko Sinapov, Matthew E. Taylor, and Peter Stone. Curriculum learning for reinforcement learning domains: A framework and survey. Journal of Machine Learning Research, 21(181):1–50, 2020.

    [7] Rémy Portelas, Cédric Colas, Lilian Weng, Katja Hofmann, and Pierre-Yves Oudeyer. Automatic curriculum learning for deep RL: A short survey. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), pages 4819–4825, 2020. doi: 10.24963/ijcai.2020/671.

    [8] Rodrigo Toro Icarte, Toryn Q. Klassen, Richard Valenzano, and Sheila A. McIlraith. Reward machines: Exploiting reward function structure in reinforcement learning. Journal of Artificial Intelligence Research, 73:1–68, 2022.

    [9] Andrew Y. Ng, Daishi Harada, and Stuart J. Russell. Policy invariance under reward transformations: Theory and application to reward shaping. In Proceedings of the International Conference on Machine Learning (ICML), pages 278–287, 1999.

    [10] Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015.

    [11] Hado van Hasselt, Arthur Guez, and David Silver. Deep reinforcement learning with double Q-learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 30, 2016.

    [12] Anna Harutyunyan, Sam Devlin, Peter Vrancx, and Ann Nowé. Expressing arbitrary reward functions as potential-based advice. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 29, 2015.

    [13] Sam Devlin and Daniel Kudenko. Dynamic potential-based reward shaping. In Proceedings of the International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS), pages 433–440, 2012.

    [14] Eric Wiewiora, Garrison W. Cottrell, and Charles Elkan. Principled methods for advising reinforcement learning agents. In Proceedings of the International Conference on Machine Learning (ICML), pages 792–799, 2003.

    [15] Grant C. Forbes, Nitish Gupta, Leonardo Villalobos-Arias, Colin M. Potts, Arnav Jhala, and David L. Roberts. Potential-based reward shaping for intrinsic motivation. In Proceedings of the International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS), pages 589–597, 2024.

    [16] Henrik Müller and Daniel Kudenko. Improving the effectiveness of potential-based reward shaping in reinforcement learning. In Proceedings of the International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS), pages 2684–2686, 2025. doi: 10.5555/3709347.3743978. Extended Abstract.

    [17] Mingxuan Li, Junzhe Zhang, and Elias Bareinboim. Automatic reward shaping from confounded offline data. In Proceedings of the 42nd International Conference on Machine Learning, volume 267, pages 36765–36793. PMLR, 2025.

    [18] Taeyoung Kim, Taemin Kang, Haechan Jeong, and Dongsoo Har. Clustering-based failed goal aware hindsight experience replay. PeerJ Computer Science, 10, 2024. doi: 10.7717/peerj-cs.2588.

    [19] Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, Vikash Kumar, and Wojciech Zaremba. Multi-goal reinforcement learning: Challenging robotics environments and request for research. arXiv preprint arXiv:1802.09464, 2018.

    [20] Rati Devidze, Parameswaran Kamalaruban, and Adish Singla. Exploration-guided reward shaping for reinforcement learning under sparse rewards. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, 2022.

    [21] José A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Unterthiner, Johannes Brandstetter, and Sepp Hochreiter. RUDDER: Return decomposition for delayed rewards. In Advances in Neural Information Processing Systems (NeurIPS), volume 32, 2019.

    [22] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.

    [23] Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. Prioritized experience replay. In Proceedings of the International Conference on Learning Representations (ICLR), 2016.

    [24] Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, and David Silver. Rainbow: Combining improvements in deep reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018.

    [25] Pierre-Luc Bacon, Jean Harb, and Doina Precup. The option-critic architecture. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 31, 2017.

    [26] Marwa Abdulhai, Dong-Ki Kim, Matthew Riemer, Miao Liu, Gerald Tesauro, and Jonathan P. How. CRADOL: Context-specific representation abstraction for deep option learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, 2022.

    [27] Shanchuan Wan, Yujin Tang, Yingtao Tian, and Tomoyuki Kaneko. DEIR: Efficient and robust exploration through discriminative-model-based episodic intrinsic rewards. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), pages 4319–4327, 2023.

    [28] Daniel Yang, Davin Tjia, Jacob Berg, Dima Damen, Pulkit Agrawal, and Abhishek Gupta. Rank2reward: Learning shaped reward functions from passive video. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), 2024.

    [29] Vishnu Sarukkai, Brennan Shacklett, Zander Majercik, Kush Bhatia, Christopher Ré, and Kayvon Fatahalian. Automated rewards via LLM-generated progress functions. arXiv preprint arXiv:2410.09187, 2024.

    下載圖示
    校外:立即公開
    QR CODE