| 研究生: |
龎懋翔 Pang, Mao-Hsiang |
|---|---|
| 論文名稱: |
透過虛實整合之目標導引強化學習於自主導航 Goal-Guided Reinforcement Learning with Sim-to-Real Transfer for Autonomous Navigation |
| 指導教授: |
連震杰
Lien, Jenn-Jier |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 117 |
| 中文關鍵詞: | 虛實整合 、模仿學習 、強化學習 、光達 、視覺 、自主導航 |
| 外文關鍵詞: | Sim-to-Real, Imitation Learning, Reinforcement Learning, LiDAR, Vision, Autonomous Navigation |
| 相關次數: | 點閱:31 下載:2 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究基於ROS2機器人系統架構,建構三種不同之局部路徑規劃(Local Planner)導航方法,包括: LiDAR-based 結合Twin Delayed Deep Deterministic Policy Gradient(TD3)演算法、 Vision-based結合Goal-Guided Transformer Soft Actor-Critic(GoT-SAC)演算法,以及 Vision-based結合Goal-Guided Transformer Soft Actor-Critic(GoT-SAC)與Imitation Learning於鍵盤和遙控器的組合之混合式導航架構。研究中將三種方法整合至地圖式導航系統,利用全域路徑規劃器提供導航目標資訊,使移動機器人能於室內環境中執行自主多點導航任務。本研究建立室內實驗場域,設計多組導航測試情境,包括多目標點導航、靜態障礙物避障、動態人員避讓及不同定位點導航測試。實驗過程量測導航成功率、碰撞次數以及目標點定位誤差等指標,以評估不同感測器資訊與學習策略對導航效能之影響。同時分析模仿學習導入後對策略收斂速度、導航穩定性及環境適應能力之改善效果。研究結果顯示,三種導航架構皆能完成室內多點導航任務,但在導航效率、避障能力及定位精度方面存在差異。LiDAR-based 結合 TD3架構於導航穩定性與定位精度方面表現良好;Vision-based 結合 GoT-SAC架構具備較佳的環境語意感知能力;而導入模仿學習後之Vision-based 結合 GoT-SAC 架構則有效提升導航成功率與策略收斂效率,並改善動態環境中的避障表現。研究成果可作為未來自主移動機器人於感測器選擇、局部路徑規劃設計及深度強化學習導航系統開發之參考依據。
This study presents three learning-based local planner frameworks for map-based autonomous mobile robot navigation under the Robot Operating System 2 (ROS2) architecture, including a LiDAR-based navigation framework using Twin Delayed Deep Deterministic Policy Gradient (TD3), an RGB Vision-based navigation framework using Goal-Guided Transformer Soft Actor-Critic (GoT-SAC), and a hybrid navigation framework integrating GoT-SAC with Imitation Learning (IL). A global planner is adopted to provide navigation goals, while the proposed local planners generate motion commands for autonomous navigation in indoor environments. The proposed methods are evaluated through multiple navigation scenarios, including multi-goal navigation, static obstacle avoidance, dynamic pedestrian avoidance, and different initial robot positions. Navigation performance is assessed using success rate, collision count, and localization error. Experimental results demonstrate that all three frameworks successfully accomplish autonomous navigation tasks with different characteristics. The LiDAR-based TD3 framework provides superior navigation stability and localization accuracy, the Vision-based with GoT-SAC framework exhibits enhanced environmental perception capability, and the proposed Vision-based with GoT-SAC and IL framework further improves navigation success rate, accelerates policy convergence, and achieves better obstacle avoidance performance in dynamic environments. The proposed frameworks provide useful references for the development of learning-based local planners for autonomous mobile robot navigation.
[1] D.S. Chaplot, D. Gandhi, A. Gupta, and R. Salakhutdinov, “Object Goal Navigation using Goal-Oriented Semantic Exploration,” Neural Information Processing Systems, Vol. 33, pp. 4247–4258, Jul. 2020.
[2] D.S. Chaplot, D. Gandhi, S. Gupta, A. Gupta, and R. Salakhutdinov, “Learning to Explore using Active Neural SLAM,” International Conference on Learning Representations, Apr. 2020.
[3] J. Chen, K. Wu, M. Hu, P.N. Suganthan, and A. Makur, “LiDAR-based End-to-end Active SLAM using Deep Reinforcement Learning in Large-scale Environments,” IEEE Transactions on Vehicular Technology, Vol. 73, no. 10, pp. 14187–14200, Oct. 2024.
[4] R. Cimurs, I.H. Suh, and J.H. Lee, “Goal-driven Autonomous Exploration through Deep Reinforcement Learning,” IEEE Robotics and Automation Letters, Vol. 7, no. 2, pp. 730–737, Apr. 2022.
[5] A. Devo, G. Mezzetti, G. Costante, M.L. Fravolini, and P. Valigi, “Towards Generalization in Target-Driven Visual Navigation by Using Deep Reinforcement Learning,” IEEE Transactions on Robotics, Vol. 36, no. 5, pp. 1546–1561, Oct. 2020.
[6] Y. Gao, J. Wu, X. Yang, and Z. Ji, “Efficient Hierarchical Reinforcement Learning for Mapless Navigation with Predictive Neighbouring Space Scoring,” IEEE Transactions on Automation Science and Engineering, Vol. 21, no. 4, pp. 5457–5472, Oct. 2024.
[7] Y. Hu, S. Wang, Y. Xie, S. Zheng, P. Shi, I. Rudas, and X. Cheng, “Deep Reinforcement Learning-based Mapless Navigation for Mobile Robot in Unknown Environment with Local Optima,” IEEE Robotics and Automation Letters, Vol. 10, no. 1, pp. 628–635, Jan. 2025.
[8] W. Huang, Y. Zhou, X. He, and C. Lv, “Goal-Guided Transformer-Enabled Reinforcement Learning for Efficient Autonomous Navigation,” IEEE Transactions on Intelligent Transportation Systems, Vol. 25, no. 2, pp. 1832–1845, Feb. 2024.
[9] Y. Kolomeytsev and D. Golembiovsky, “Hybrid Motion Planning with Deep Reinforcement Learning for Mobile Robot Navigation,” ArXiv, Dec. 2025.
[10] M.F.R. Lee and S.H. Yusuf, “Mobile Robot Navigation Using Deep Reinforcement Learning,” Processes, Vol. 10, no. 12, pp. 2748–2770, Nov. 2022.
[11] V.R.F. Miranda, A.A. Neto, G.M. Freitas, and L.A. Mozelli, “Generalization in Deep Reinforcement Learning for Robotic Navigation by Reward Shaping,” IEEE Transactions on Industrial Electronics, Vol. 71, no. 6, pp. 6013–6020, Jun. 2024.
[12] M. Pfeiffer, S. Shukla, M. Turchetta, C. Cadena, A. Krause, and R. Siegwart, “Reinforced Imitation: Sample Efficient Deep Reinforcement Learning for Mapless Navigation by Leveraging Prior Demonstrations,” IEEE Robotics and Automation Letters, Vol. 3, no. 4, pp. 4423–4430, Oct. 2018.
[13] M. Pfeiffer, M. Schaeuble, J. Nieto, R. Siegwart, and C. Cadena, “From Perception to Decision: A Data-driven Approach to End-to-end Motion Planning for Autonomous Ground Robots,” IEEE International Conference on Robotics and Automation, pp. 1527–1533, May. 2018.
[14] N. Selvaraj, A. Mitrevski, and S. Houben, “Are Learning-Based Approaches Ready for Real-World Indoor Navigation? A Case for Imitation Learning,” European Conference on Mobile Robots, Padova, Italy, pp. 1–7, Aug. 2025.
[15] V.D. Sharma, J. Lee, M. Andrews, and I. Hadžić, “Hybrid Classical/RL Local Planner for Ground Robot Navigation,” ArXiv, Oct. 2024.
[16] Y. Tang, C. Zhao, J. Wang, C. Zhang, Q. Sun, and W.X. Zheng, “Perception and Navigation in Autonomous Systems in the Era of Learning: A survey,” IEEE Trans. Neural Netw. Learn. Syst., early access, Apr. 2022.
[17] T. Wen, X. Wang, Z. Zheng, and Z. Sun, “A DRL-based Path Planning Method for Wheeled Mobile Robots in Unknown Environments,” Computers and Electrical Engineering, Vol. 118, Sep. 2024.
[18] X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, and Y.F. Wang, “Vision-language Navigation Policy Learning and Adaptation,” IEEE Trans. Pattern Anal. Mach. Intell., Vol. 43, no. 12, pp. 4205–4216, Dec. 2021.
[19] Y. Wu, Y. Li, W. Li, H. Li, and R. Lu, “Robust Lidar-based Localization Scheme for Unmanned Ground Vehicle via Multisensor Fusion,” IEEE Trans. Neural Netw. Learn. Syst., Vol. 32, no. 12, pp. 5633–5643, Dec. 2021.
[20] Y. Yin, Z. Chen, G. Liu, J. Yin, and J. Guo, “Autonomous Navigation of Mobile Robots in Unknown Environments using Off-policy Reinforcement Learning with Curriculum Learning,” Expert Systems with Applications, Vol. 274, Aug. 2024.
[21] A. Zeng, S. Song, S. Welker, J. Lee, A. Rodriguez, and T. Funkhouser, “Learning Synergies between Pushing and Grasping with Self-supervised Deep Reinforcement Learning,” Intelligent Robots and Systems, Vol. 32, pp. 1069–1075, Mar. 2018.