簡易檢索 / 詳目顯示

研究生: 龎懋翔
Pang, Mao-Hsiang
論文名稱: 透過虛實整合之目標導引強化學習於自主導航
Goal-Guided Reinforcement Learning with Sim-to-Real Transfer for Autonomous Navigation
指導教授: 連震杰
Lien, Jenn-Jier
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 117
中文關鍵詞: 虛實整合模仿學習強化學習光達視覺自主導航
外文關鍵詞: Sim-to-Real, Imitation Learning, Reinforcement Learning, LiDAR, Vision, Autonomous Navigation
相關次數: 點閱:31下載:2
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究基於ROS2機器人系統架構,建構三種不同之局部路徑規劃(Local Planner)導航方法,包括: LiDAR-based 結合Twin Delayed Deep Deterministic Policy Gradient(TD3)演算法、 Vision-based結合Goal-Guided Transformer Soft Actor-Critic(GoT-SAC)演算法,以及 Vision-based結合Goal-Guided Transformer Soft Actor-Critic(GoT-SAC)與Imitation Learning於鍵盤和遙控器的組合之混合式導航架構。研究中將三種方法整合至地圖式導航系統,利用全域路徑規劃器提供導航目標資訊,使移動機器人能於室內環境中執行自主多點導航任務。本研究建立室內實驗場域,設計多組導航測試情境,包括多目標點導航、靜態障礙物避障、動態人員避讓及不同定位點導航測試。實驗過程量測導航成功率、碰撞次數以及目標點定位誤差等指標,以評估不同感測器資訊與學習策略對導航效能之影響。同時分析模仿學習導入後對策略收斂速度、導航穩定性及環境適應能力之改善效果。研究結果顯示,三種導航架構皆能完成室內多點導航任務,但在導航效率、避障能力及定位精度方面存在差異。LiDAR-based 結合 TD3架構於導航穩定性與定位精度方面表現良好;Vision-based 結合 GoT-SAC架構具備較佳的環境語意感知能力;而導入模仿學習後之Vision-based 結合 GoT-SAC 架構則有效提升導航成功率與策略收斂效率,並改善動態環境中的避障表現。研究成果可作為未來自主移動機器人於感測器選擇、局部路徑規劃設計及深度強化學習導航系統開發之參考依據。

    This study presents three learning-based local planner frameworks for map-based autonomous mobile robot navigation under the Robot Operating System 2 (ROS2) architecture, including a LiDAR-based navigation framework using Twin Delayed Deep Deterministic Policy Gradient (TD3), an RGB Vision-based navigation framework using Goal-Guided Transformer Soft Actor-Critic (GoT-SAC), and a hybrid navigation framework integrating GoT-SAC with Imitation Learning (IL). A global planner is adopted to provide navigation goals, while the proposed local planners generate motion commands for autonomous navigation in indoor environments. The proposed methods are evaluated through multiple navigation scenarios, including multi-goal navigation, static obstacle avoidance, dynamic pedestrian avoidance, and different initial robot positions. Navigation performance is assessed using success rate, collision count, and localization error. Experimental results demonstrate that all three frameworks successfully accomplish autonomous navigation tasks with different characteristics. The LiDAR-based TD3 framework provides superior navigation stability and localization accuracy, the Vision-based with GoT-SAC framework exhibits enhanced environmental perception capability, and the proposed Vision-based with GoT-SAC and IL framework further improves navigation success rate, accelerates policy convergence, and achieves better obstacle avoidance performance in dynamic environments. The proposed frameworks provide useful references for the development of learning-based local planners for autonomous mobile robot navigation.

    摘要 I Abstract II 誌謝 III List of Tables VII List of Figures VIII Chapter 1 Introduction 1 1.1 Research Motivation and Objectives 1 1.2 Global Framework 4 1.3 Related Work 15 1.4 Contribution 18 Chapter 2 System Setup and Specification 20 2.1 System Setup 20 2.2 Hardware Specifications 25 Chapter 3 LiDAR-based Reinforcement Learning Local Path Planning 30 3.1 LiDAR-based Reinforcement Learning Local Path Planning 30 3.1.0 LiDAR-based Reinforcement Learning Local Path Planning Framework 31 3.1.1 LiDAR-based Reinforcement Learning Local Path Planning Training Framework 31 3.1.2 LiDAR-based Reinforcement Learning Local Path Planning Inference Framework 43 3.1.3 Loss Function 47 3.2 Heuristic Function 50 Chapter 4 Vision-based Reinforcement Learning Local Path Planning 54 4.1 Vision-based Reinforcement Learning Local Path Planning 54 4.1.0 Motivation of Vision-based Reinforcement Learning Local Path Planning 55 4.1.1 Vision-based Reinforcement Learning Local Path Planning Training Framework 57 4.1.2 Vision-based Reinforcement Learning Local Path Planning Inference Framework 65 4.1.3 Loss Function 68 4.2 Sim-to-Real Reinforcement Learning and Imitation Learning Training Framework 72 Chapter 5 Experimental Result 76 5.1 Data Collection 76 5.2 Metrics 79 5.3 Experimental Results –LiDAR-based RL Results Analysis 84 5.4 Experimental Results –Vision-based RL-IL Results Analysis 89 5.5 Experimental Results –Vision-based RL-IL Fine-tuning Amount Analysis 95 Chapter 6 Conclusions and Future Work 99 6.1 Conclusions 99 6.2 Future Work 101 Reference 104

    [1] D.S. Chaplot, D. Gandhi, A. Gupta, and R. Salakhutdinov, “Object Goal Navigation using Goal-Oriented Semantic Exploration,” Neural Information Processing Systems, Vol. 33, pp. 4247–4258, Jul. 2020.
    [2] D.S. Chaplot, D. Gandhi, S. Gupta, A. Gupta, and R. Salakhutdinov, “Learning to Explore using Active Neural SLAM,” International Conference on Learning Representations, Apr. 2020.
    [3] J. Chen, K. Wu, M. Hu, P.N. Suganthan, and A. Makur, “LiDAR-based End-to-end Active SLAM using Deep Reinforcement Learning in Large-scale Environments,” IEEE Transactions on Vehicular Technology, Vol. 73, no. 10, pp. 14187–14200, Oct. 2024.
    [4] R. Cimurs, I.H. Suh, and J.H. Lee, “Goal-driven Autonomous Exploration through Deep Reinforcement Learning,” IEEE Robotics and Automation Letters, Vol. 7, no. 2, pp. 730–737, Apr. 2022.
    [5] A. Devo, G. Mezzetti, G. Costante, M.L. Fravolini, and P. Valigi, “Towards Generalization in Target-Driven Visual Navigation by Using Deep Reinforcement Learning,” IEEE Transactions on Robotics, Vol. 36, no. 5, pp. 1546–1561, Oct. 2020.
    [6] Y. Gao, J. Wu, X. Yang, and Z. Ji, “Efficient Hierarchical Reinforcement Learning for Mapless Navigation with Predictive Neighbouring Space Scoring,” IEEE Transactions on Automation Science and Engineering, Vol. 21, no. 4, pp. 5457–5472, Oct. 2024.
    [7] Y. Hu, S. Wang, Y. Xie, S. Zheng, P. Shi, I. Rudas, and X. Cheng, “Deep Reinforcement Learning-based Mapless Navigation for Mobile Robot in Unknown Environment with Local Optima,” IEEE Robotics and Automation Letters, Vol. 10, no. 1, pp. 628–635, Jan. 2025.
    [8] W. Huang, Y. Zhou, X. He, and C. Lv, “Goal-Guided Transformer-Enabled Reinforcement Learning for Efficient Autonomous Navigation,” IEEE Transactions on Intelligent Transportation Systems, Vol. 25, no. 2, pp. 1832–1845, Feb. 2024.
    [9] Y. Kolomeytsev and D. Golembiovsky, “Hybrid Motion Planning with Deep Reinforcement Learning for Mobile Robot Navigation,” ArXiv, Dec. 2025.
    [10] M.F.R. Lee and S.H. Yusuf, “Mobile Robot Navigation Using Deep Reinforcement Learning,” Processes, Vol. 10, no. 12, pp. 2748–2770, Nov. 2022.
    [11] V.R.F. Miranda, A.A. Neto, G.M. Freitas, and L.A. Mozelli, “Generalization in Deep Reinforcement Learning for Robotic Navigation by Reward Shaping,” IEEE Transactions on Industrial Electronics, Vol. 71, no. 6, pp. 6013–6020, Jun. 2024.
    [12] M. Pfeiffer, S. Shukla, M. Turchetta, C. Cadena, A. Krause, and R. Siegwart, “Reinforced Imitation: Sample Efficient Deep Reinforcement Learning for Mapless Navigation by Leveraging Prior Demonstrations,” IEEE Robotics and Automation Letters, Vol. 3, no. 4, pp. 4423–4430, Oct. 2018.
    [13] M. Pfeiffer, M. Schaeuble, J. Nieto, R. Siegwart, and C. Cadena, “From Perception to Decision: A Data-driven Approach to End-to-end Motion Planning for Autonomous Ground Robots,” IEEE International Conference on Robotics and Automation, pp. 1527–1533, May. 2018.
    [14] N. Selvaraj, A. Mitrevski, and S. Houben, “Are Learning-Based Approaches Ready for Real-World Indoor Navigation? A Case for Imitation Learning,” European Conference on Mobile Robots, Padova, Italy, pp. 1–7, Aug. 2025.
    [15] V.D. Sharma, J. Lee, M. Andrews, and I. Hadžić, “Hybrid Classical/RL Local Planner for Ground Robot Navigation,” ArXiv, Oct. 2024.
    [16] Y. Tang, C. Zhao, J. Wang, C. Zhang, Q. Sun, and W.X. Zheng, “Perception and Navigation in Autonomous Systems in the Era of Learning: A survey,” IEEE Trans. Neural Netw. Learn. Syst., early access, Apr. 2022.
    [17] T. Wen, X. Wang, Z. Zheng, and Z. Sun, “A DRL-based Path Planning Method for Wheeled Mobile Robots in Unknown Environments,” Computers and Electrical Engineering, Vol. 118, Sep. 2024.
    [18] X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, and Y.F. Wang, “Vision-language Navigation Policy Learning and Adaptation,” IEEE Trans. Pattern Anal. Mach. Intell., Vol. 43, no. 12, pp. 4205–4216, Dec. 2021.
    [19] Y. Wu, Y. Li, W. Li, H. Li, and R. Lu, “Robust Lidar-based Localization Scheme for Unmanned Ground Vehicle via Multisensor Fusion,” IEEE Trans. Neural Netw. Learn. Syst., Vol. 32, no. 12, pp. 5633–5643, Dec. 2021.
    [20] Y. Yin, Z. Chen, G. Liu, J. Yin, and J. Guo, “Autonomous Navigation of Mobile Robots in Unknown Environments using Off-policy Reinforcement Learning with Curriculum Learning,” Expert Systems with Applications, Vol. 274, Aug. 2024.
    [21] A. Zeng, S. Song, S. Welker, J. Lee, A. Rodriguez, and T. Funkhouser, “Learning Synergies between Pushing and Grasping with Self-supervised Deep Reinforcement Learning,” Intelligent Robots and Systems, Vol. 32, pp. 1069–1075, Mar. 2018.

    下載圖示
    校外:立即公開
    QR CODE