| 研究生: |
沈孔堯 Shen, Kung-Yao |
|---|---|
| 論文名稱: |
雙足機器人步態控制:基於規則之資料生成與監督式預訓練 Rule-Based Data Generation and Supervised Pre-training for Bipedal Locomotion Control |
| 指導教授: |
蘇文鈺
Su, Wen-Yu |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 醫學資訊研究所 Institute of Medical Informatics |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 96 |
| 中文關鍵詞: | 雙足機器人 、模仿學習 、行為複製 、步態控制 |
| 外文關鍵詞: | Bipedal Robot, Imitation Learning, Behavior Cloning, Locomotion Control |
| 相關次數: | 點閱:106 下載:2 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本論文旨在改善雙足機器人模型步態控制中的訓練效率,並提出一套結合仿生控制,監督式學習到強化式學習的混合流程。我們首先實作一套控制規則,整合人類步行的相位切換與重心轉移機制,利用軌跡規劃與正逆向運動學,建立可解釋且具物理約束的規則控制器。接著在 Isaac Sim 模擬環境中,收集機器人狀態與規則控制指令,並濾除失敗與不穩定樣本,提升資料可用性。最後,利用前述資料,我們訓練出一個多層感知器模型(MLP)。此模型旨在複製穩定的步態,為後續強化式學習提供初始模型權重,避免從零開始的未知探索。
在相同的初始化與終止條件下,我們比較了規則控制器與 MLP 模型的表現。連續採樣 10 個完整步態週期的結果為,MLP 模型開始行走至跌倒的穩定步數平均為 20.3 步(中位數 23 步); 規則控制器的穩定步數平均為 10.6 步(中位數 12 步)。結果顯示,MLP 模型的表現優於規則控制器,我們推測這主要得益於失敗樣本過濾機制,使 MLP 模型能專注於學習成功的範例資料。另一方面,失敗分析指出,模型控制系統效果仍受隨機擾動與模型預測分布所影響,未來將採用更具泛化性的強化式學習以解決此問題。基於上述,本論文完成了從零到一的步態資料生成與初步模型學習驗證,也為後續導入強化式學習,提供了預訓練初始模型權重,預期將可縮減傳統強化式學習初期的探索週期,提升模型訓練效率。
In this thesis, we aim to improve training efficiency in bipedal robot gait control by proposing a hybrid framework that integrates biomimetic control, supervised learning, and reinforcement learning (RL). First, we implement a rule-based controller that incorporates the phase transition and center-of-mass shifting mechanisms of human walking. By utilizing trajectory planning and forward and inverse kinematics, we construct an interpretable and physically constrained rule-based controller. Next, in Isaac Sim, we collect robot states and rule-based control commands, and filter out failed and unstable samples to enhance data usability. Finally, using the aforementioned data, we train a Multi-Layer Perceptron (MLP) model. This model is designed to replicate stable gaits and provide initial model weights for subsequent RL, thereby avoiding undirected, from-scratch exploration.
Under the same initialization and termination conditions, we compared the performance of the rule-based controller and the MLP model. Based on consecutive sampling across 10 complete gait cycles, the MLP model achieved an average of 20.3 stable steps (median: 23 steps) from the start of walking to falling, whereas the rule-based controller averaged 10.6 stable steps (median: 12 steps). The results show that the MLP model outperforms the rule-based controller. We hypothesize that this is primarily due to the failure sample filtering mechanism, which enables the MLP model to focus on learning from successful demonstrations. On the other hand, failure analysis indicates that the control system's performance is still affected by random perturbations and the model's prediction distribution. To address this, a more generalizable RL approach will be adopted in the future. Based on the above, this thesis accomplishes the "zero-to-one" generation of gait data and initial model learning verification. Furthermore, it provides pre-trained initial model weights for the subsequent introduction of RL, which is expected to shorten the initial exploration period of traditional RL and improve model training efficiency.
[1] B.D. Argall, S. Chernova, M. Veloso, B. Browning, A survey of robot learning from demonstration, Robotics and Autonomous Systems, 57, 5, 469–483, 2009.
[2] A. Aristidou, J. Lasenby, Inverse Kinematics: a review of existing techniques and introduction of a new fast iterative solver, 2009.
[3] K. Arulkumaran, M.P. Deisenroth, M. Brundage, A.A. Bharath, Deep Reinforcement Learning: A Brief Survey, IEEE Signal Processing Magazine, 34, 6, 26–38, 2017.
[4] Autodesk, Inc., Fusion 360, https://www.autodesk.com/tw/products/fusion-360/overview, 2026.
[5] L. Bao, J. Humphreys, T. Peng, C. Zhou, Deep reinforcement learning for robotic bipedal locomotion: A brief survey, Artificial Intelligence Review, 59, 1, 38, 2025.
[6] J. Carius, F. Farshidian, M. Hutter, MPC-Net: A First Principles Guided Policy Search, IEEE Robotics and Automation Letters (RA-L), 5, 2, 2897–2904, 2020.
[7] X. Cheng, Y. Ji, J. Chen, R. Yang, G. Yang, X. Wang, Expressive whole-body control for humanoid robots, Arxiv Preprint Arxiv:2402.16796, 2024.
[8] J. Dao, H. Duan, A. Fern, Sim-to-Real Learning for Humanoid Box Loco-Manipulation, 2024 IEEE International Conference on Robotics and Automation (ICRA), 16930–16936, 2024.
[9] Docker Inc., Docker Documentation, https://docs.docker.com/, 2026.
[10] G.F. Franklin, J.D. Powell, M.L. Workman, others, Digital control of dynamic systems, Addison-Wesley, Menlo Park, CA, 1998.
[11] X. Gu, Y.-J. Wang, J. Chen, Humanoid-Gym: Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real Transfer, 2024.
[12] J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, M. Hutter, Learning agile and dynamic motor skills for legged robots, Science Robotics, 4, 26, eaau5872, 2019.
[13] A. Irpan, Deep Reinforcement Learning Doesn't Work Yet, https://www.alexirpan.com/2018/02/14/rl-hard.html, 2018.
[14] T. Johannink, S. Bahl, A. Nair, J. Luo, A. Kumar, M. Loskyll, J.A. Ojea, E. Solowjow, S. Levine, Residual Reinforcement Learning for Robot Control, 2019 International Conference on Robotics and Automation (ICRA), 6023–6029, 2019.
[15] L.P. Kaelbling, M.L. Littman, A.W. Moore, Reinforcement learning: A survey, Journal of Artificial Intelligence Research, 4, 237–285, 1996.
[16] S. Kajita, H. Hirukawa, K. Harada, K. Yokoi, Introduction to Humanoid Robotics, Springer, Berlin, Heidelberg, 2014.
[17] S. Kajita, F. Kanehiro, K. Kaneko, K. Fujiwara, K. Harada, K. Yokoi, H. Hirukawa, Biped walking pattern generation by using preview control of zero-moment point, 2003 IEEE International Conference on Robotics and Automation (Cat. No.03ch37422), 2, 1620–1626, 2003.
[18] S. Kajita, M. Morisawa, K. Miura, S. Nakaoka, K. Harada, K. Kaneko, F. Kanehiro, K. Yokoi, Biped walking stabilization based on linear inverted pendulum tracking, Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 4489–4496, 2010.
[19] R.C. Luo, C.-C. Chen, Biped Walking Trajectory Generator Based on Three-Mass With Angular Momentum Model Using Model Predictive Control, IEEE Transactions on Industrial Electronics, 63, 1, 268–276, 2016.
[20] K. Lynch, F. Park, Modern Robotics: Mechanics, Planning, and Control, Cambridge University Press, Cambridge, UK, 2017.
[21] S. Macenski, T. Foote, B. Gerkey, C. Lalancette, W. Woodall, Robot Operating System 2: Design, architecture, and uses in the wild, Science Robotics, 7, 66, eabm6074, 2022.
[22] V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, G. State, Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning, 2021.
[23] H.T. Najm, A. Sabah Al-Araji, N.S. Ahmad, Mobile and Humanoid Robots in Healthcare: A Systematic Review of Real-World Applications and Deployment Challenges, IEEE Access, 14, 77677–77700, 2026.
[24] NVIDIA Corporation, NVIDIA PhysX SDK 5: A Scalable Multi-Platform Physics Solution, https://github.com/NVIDIA-Omniverse/PhysX, 2025.
[25] NVIDIA Corporation, ROS & ROS2 Bridge - Isaac Sim 4.2.0 Documentation, https://docs.isaacsim.omniverse.nvidia.com/4.2.0/features/external_communication/ext_omni_isaac_ros_bridge.html, 2025.
[26] NVIDIA Corporation, Tuning Joint Drive Gains - Isaac Sim 4.5.0 Documentation, https://docs.isaacsim.omniverse.nvidia.com/4.5.0/robot_setup/joint_tuning.html, 2025.
[27] NVIDIA Corporation, NVIDIA Isaac Sim: Robotics Simulation and Synthetic Data Generation, https://developer.nvidia.com/isaac/sim, 2026.
[28] NVIDIA Corporation, NVIDIA Isaac Lab, https://developer.nvidia.com/isaac/lab, 2026.
[29] NVIDIA Corporation, Livestream Clients - Isaac Sim 5.0.0 Documentation, https://docs.isaacsim.omniverse.nvidia.com/5.0.0/installation/manual_livestream_clients.html, 2026.
[30] Open Robotics, ROS 2 Documentation: URDF Tutorials (Humble), https://docs.ros.org/en/humble/Tutorials/Intermediate/URDF/URDF-Main.html, 2026.
[31] T. Osa, J. Pajarinen, G. Neumann, J.A. Bagnell, P. Abbeel, J. Peters, An Algorithmic Perspective on Imitation Learning, Foundations and Trends in Robotics, 7, 1–2, 1–179, 2018.
[32] X.B. Peng, Z. Ma, P. Abbeel, S. Levine, A. Kanazawa, AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control, ACM Transactions on Graphics (TOG), 40, 4, 20, 2021.
[33] M.-C. Popescu, V.E. Balas, L. Perescu-Popescu, N. Mastorakis, Multilayer perceptron and neural networks, WSEAS Transactions on Circuits and Systems, 8, 7, 579–588, 2009.
[34] M. Posa, C. Cantu, R. Tedrake, A direct method for trajectory optimization of rigid bodies through contact, The International Journal of Robotics Research, 33, 1, 69–81, 2014.
[35] A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, S. Levine, Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations, 2018.
[36] N. Rudin, D. Hoeller, P. Reist, M. Hutter, Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning, Proceedings of the 5th Conference on Robot Learning, 164, 91–100, 2022.
[37] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal Policy Optimization Algorithms, Arxiv Preprint Arxiv:1707.06347, 2017.
[38] screamlab, fusion2urdf-ros2: Fusion360 to URDF converter for ROS2, https://github.com/screamlab/fusion2urdf-ros2, 2026.
[39] screamlab, fusion_xacro2urdf2unity, https://github.com/screamlab/fusion_xacro2urdf2unity, 2026.
[40] J. Siekmann, K. Green, J. Warila, A. Fern, J. Hurst, Blind Bipedal Stair Traversal via Sim-to-Real Reinforcement Learning, Robotics: Science and Systems, 2021.
[41] J. Tan, T. Zhang, E. Coumans, A. Iscen, Y. Bai, D. Hafner, S. Bohez, V. Vanhoucke, Sim-to-Real: Learning Agile Locomotion For Quadruped Robots, Proceedings of Robotics: Science and Systems, 2018.
[42] G. Tevatia, S. Schaal, Inverse kinematics for humanoid robots, Proceedings 2000 ICRA. Millennium Conference. IEEE International Conference on Robotics and Automation. Symposia Proceedings (Cat. No.00ch37065), 1, 294–299, 2000.
[43] X. Wen, L. Wang, Y. Tao, H. Lai, H. Liu, Reinforcement Learning-Based Adaptive Motion Control of Humanoid Robots on Multi-Terrain, Applied Sciences, 16, 5, 2371, 2026.
[44] P.-B. Wieber, Trajectory Free Linear Model Predictive Control for Stable Walking in the Presence of Strong Perturbations, Proceedings of the 6th IEEE-RAS International Conference on Humanoid Robots (Humanoids), 137–142, 2006.
[45] K. Yin, K. Loken, M. van de Panne, SIMBICON: Simple Biped Locomotion Control, ACM Transactions on Graphics (TOG), 26, 3, 105, 2007.
[46] T. Zhang, B. Zheng, R. Nai, Y. Hu, Y.-J. Wang, G. Chen, F. Lin, J. Li, C. Hong, K. Sreenath, others, Hub: Learning extreme humanoid balance, Arxiv Preprint Arxiv:2505.07294, 2025.
[47] W. Zhao, J.P. Queralta, T. Westerlund, Sim-to-Real Transfer in Deep Reinforcement Learning for Robotics: a Survey, 2020 IEEE Symposium Series on Computational Intelligence (SSCI), 737–744, 2020.