簡易檢索 / 詳目顯示

研究生: 沈孔堯
Shen, Kung-Yao
論文名稱: 雙足機器人步態控制:基於規則之資料生成與監督式預訓練
Rule-Based Data Generation and Supervised Pre-training for Bipedal Locomotion Control
指導教授: 蘇文鈺
Su, Wen-Yu
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 醫學資訊研究所
Institute of Medical Informatics
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 96
中文關鍵詞: 雙足機器人 、模仿學習 、行為複製 、步態控制
外文關鍵詞: Bipedal Robot, Imitation Learning, Behavior Cloning, Locomotion Control
相關次數: 點閱:106  下載:2 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本論文旨在改善雙足機器人模型步態控制中的訓練效率,並提出一套結合仿生控制,監督式學習到強化式學習的混合流程。我們首先實作一套控制規則,整合人類步行的相位切換與重心轉移機制,利用軌跡規劃與正逆向運動學,建立可解釋且具物理約束的規則控制器。接著在 Isaac Sim 模擬環境中,收集機器人狀態與規則控制指令,並濾除失敗與不穩定樣本,提升資料可用性。最後,利用前述資料,我們訓練出一個多層感知器模型(MLP)。此模型旨在複製穩定的步態,為後續強化式學習提供初始模型權重,避免從零開始的未知探索。

    在相同的初始化與終止條件下,我們比較了規則控制器與 MLP 模型的表現。連續採樣 10 個完整步態週期的結果為,MLP 模型開始行走至跌倒的穩定步數平均為 20.3 步(中位數 23 步); 規則控制器的穩定步數平均為 10.6 步(中位數 12 步)。結果顯示,MLP 模型的表現優於規則控制器,我們推測這主要得益於失敗樣本過濾機制,使 MLP 模型能專注於學習成功的範例資料。另一方面,失敗分析指出,模型控制系統效果仍受隨機擾動與模型預測分布所影響,未來將採用更具泛化性的強化式學習以解決此問題。基於上述,本論文完成了從零到一的步態資料生成與初步模型學習驗證,也為後續導入強化式學習,提供了預訓練初始模型權重,預期將可縮減傳統強化式學習初期的探索週期,提升模型訓練效率。

    In this thesis, we aim to improve training efficiency in bipedal robot gait control by proposing a hybrid framework that integrates biomimetic control, supervised learning, and reinforcement learning (RL). First, we implement a rule-based controller that incorporates the phase transition and center-of-mass shifting mechanisms of human walking. By utilizing trajectory planning and forward and inverse kinematics, we construct an interpretable and physically constrained rule-based controller. Next, in Isaac Sim, we collect robot states and rule-based control commands, and filter out failed and unstable samples to enhance data usability. Finally, using the aforementioned data, we train a Multi-Layer Perceptron (MLP) model. This model is designed to replicate stable gaits and provide initial model weights for subsequent RL, thereby avoiding undirected, from-scratch exploration.

    Under the same initialization and termination conditions, we compared the performance of the rule-based controller and the MLP model. Based on consecutive sampling across 10 complete gait cycles, the MLP model achieved an average of 20.3 stable steps (median: 23 steps) from the start of walking to falling, whereas the rule-based controller averaged 10.6 stable steps (median: 12 steps). The results show that the MLP model outperforms the rule-based controller. We hypothesize that this is primarily due to the failure sample filtering mechanism, which enables the MLP model to focus on learning from successful demonstrations. On the other hand, failure analysis indicates that the control system's performance is still affected by random perturbations and the model's prediction distribution. To address this, a more generalizable RL approach will be adopted in the future. Based on the above, this thesis accomplishes the "zero-to-one" generation of gait data and initial model learning verification. Furthermore, it provides pre-trained initial model weights for the subsequent introduction of RL, which is expected to shorten the initial exploration period of traditional RL and improve model training efficiency.

    摘要 i Abstract ii 誌謝 iii Table of Contents iv List of Tables viii List of Figures ix Chapter 1. Introduction 1 1.1. Background and Motivation 1 1.2. Research Objective and Methodology 2 1.3. Research Contributions 2 1.4. Thesis Structure 3 Chapter 2. Background and Related Work 4 2.1. Nomenclature 4 2.2. Background 5 2.2.1. Fundamentals of Bipedal Locomotion Control 5 2.2.1.1. Center of Mass and Zero Moment Point 6 2.2.1.2. Inverse Kinematics for Leg Control 7 2.2.2. Physics Simulation and ROS 2 Bridging 8 2.2.2.1. NVIDIA Isaac Sim and PhysX 8 2.2.2.2. ROS 2 Communication and the Isaac Sim Bridge 8 2.2.3. Behavior Cloning 9 2.2.4. Reinforcement Learning 10 2.3. Related Work 10 2.3.1. Model-based Bipedal Control Methods 11 2.3.1.1. Strengths and Limitations 12 2.3.2. Reinforcement Learning and Proximal Policy Optimization in Robotics 13 2.3.2.1. Critical Limitations 14 2.3.3. Integrating Expert Priors with Learning-based Control 15 2.3.3.1. Residual Learning and Hybrid Architectures 15 2.3.3.2. Behavior Cloning from Expert Systems 16 2.3.3.3. Behavior Cloning for Reinforcement Learning Initialization 16 2.3.3.4. Position of this Thesis 17 Chapter 3. Robot Modeling and System Architecture 18 3.1. Physical Modeling and Parameterization 18 3.1.1. Mechanical Design 18 3.1.2. URDF Export 19 3.1.3. Notations Relative to Biped Mechanism 20 3.2. Remote Simulation and Distributed Architecture 20 3.2.1. Containerized Deployment 21 3.2.2. ROS 2 and Isaac Sim Integration 21 3.3. Physics Parameter Tuning in Simulation 22 3.4. System-Level Representation for Control and Simulation 22 Chapter 4. Rule-Based Balance and Gait Planning System 24 4.1. High Level Phase State Machine Design 24 4.1.1. Phase Offset and Time Parameters 25 4.2. Mid Level Phase-Dependent Trajectory Synthesis 28 4.2.1. Coordinate Frames and Notation 28 4.2.2. Initialization-to-Single-Support Trajectory 30 4.2.3. Single-Support-to-Double-Support Trajectory 31 4.2.4. Double-Support-to-Single-Support Trajectory 35 4.2.5. Summary 37 4.3. Low Level Controllers 38 4.3.1. Low Level Control Architecture 38 4.3.2. Assumptions 38 4.3.3. Stance Leg Control Node 39 4.3.4. Swing Leg Control Node 39 4.3.4.1. IK Formulation and Foot-Parallel-to-Ground Constraint 39 4.3.4.2. Joint Limits, Reachability Checks, and Fallback Policy 49 4.3.5. Counterweight Control Node 50 4.3.6. Summary 53 Chapter 5. Data-Driven Pre-training with Multi-Layer Perceptrons 54 5.1. Expert Demonstrations via Rule-Based System 54 5.1.1. Rule-based controller as expert policy 54 5.1.2. Observation Space and Action Space 55 5.2. Real-Time Data Collection in Isaac Sim and ROS 2 57 5.2.1. ROS 2 Subscription and Observation Assembly 57 5.2.2. Episode validation and dirty-data rollback mechanism 58 5.2.3. Raw data storage format 58 5.3. Data Preprocessing and Feature Engineering 59 5.3.1. Observation normalization strategy 59 5.3.2. Action normalization and symmetric scaling 60 5.4. MLP Training with Behavior Cloning 61 5.4.1. Decoupled network architecture design (Body and Head) 61 5.4.2. Training process and optimization 62 5.4.3. Evaluation Metrics 62 5.5. Online Inference and RL Foundation 63 5.5.1. Real-time inference deployment 63 5.5.2. Action post-processing 63 5.5.3. Initializing future RL agents 64 Chapter 6. Experiments, Results, and Failure Analysis 65 6.1. Experimental Setup 65 6.1.1. Control parameters 65 6.1.2. Experiment Initialization and Termination Criteria 66 6.2. Rule-Based Controller: Experimental Setup and Performance 66 6.2.1. Experimental Setup 66 6.2.1.1. Control Parameters 66 6.2.2. Performance 67 6.3. MLP Policy: Experimental Setup and Performance 69 6.3.1. Experimental Setup 69 6.3.1.1. Training Data Collection parameters 69 6.3.1.2. Training Data Preprocessing parameters 69 6.3.1.3. Data Split parameters 69 6.3.1.4. MLP Training parameters 69 6.3.2. Performance 70 6.4. Analysis and Discussion 71 6.4.1. Rule-based and MLP Policy Controllers Performance Comparison 71 6.4.2. Failure discussion 74 Chapter 7. Conclusion and Future Work 76 7.1. Conclusion 76 7.1.1. Addressing the Data Bottleneck in Embodied AI 76 7.1.2. Efficient Supervised Learning, Performance Breakthrough, and Physical Interpretability 77 7.2. Future Work 77 7.2.1. Transitioning to Reinforcement Learning in Isaac Lab 77 7.2.2. Domain Randomization and Robustness for Sim-to-Real 78 Reference 79

    [1] B.D. Argall, S. Chernova, M. Veloso, B. Browning, A survey of robot learning from demonstration, Robotics and Autonomous Systems, 57, 5, 469–483, 2009.
    [2] A. Aristidou, J. Lasenby, Inverse Kinematics: a review of existing techniques and introduction of a new fast iterative solver, 2009.
    [3] K. Arulkumaran, M.P. Deisenroth, M. Brundage, A.A. Bharath, Deep Reinforcement Learning: A Brief Survey, IEEE Signal Processing Magazine, 34, 6, 26–38, 2017.
    [4] Autodesk, Inc., Fusion 360, https://www.autodesk.com/tw/products/fusion-360/overview, 2026.
    [5] L. Bao, J. Humphreys, T. Peng, C. Zhou, Deep reinforcement learning for robotic bipedal locomotion: A brief survey, Artificial Intelligence Review, 59, 1, 38, 2025.
    [6] J. Carius, F. Farshidian, M. Hutter, MPC-Net: A First Principles Guided Policy Search, IEEE Robotics and Automation Letters (RA-L), 5, 2, 2897–2904, 2020.
    [7] X. Cheng, Y. Ji, J. Chen, R. Yang, G. Yang, X. Wang, Expressive whole-body control for humanoid robots, Arxiv Preprint Arxiv:2402.16796, 2024.
    [8] J. Dao, H. Duan, A. Fern, Sim-to-Real Learning for Humanoid Box Loco-Manipulation, 2024 IEEE International Conference on Robotics and Automation (ICRA), 16930–16936, 2024.
    [9] Docker Inc., Docker Documentation, https://docs.docker.com/, 2026.
    [10] G.F. Franklin, J.D. Powell, M.L. Workman, others, Digital control of dynamic systems, Addison-Wesley, Menlo Park, CA, 1998.
    [11] X. Gu, Y.-J. Wang, J. Chen, Humanoid-Gym: Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real Transfer, 2024.
    [12] J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, M. Hutter, Learning agile and dynamic motor skills for legged robots, Science Robotics, 4, 26, eaau5872, 2019.
    [13] A. Irpan, Deep Reinforcement Learning Doesn't Work Yet, https://www.alexirpan.com/2018/02/14/rl-hard.html, 2018.
    [14] T. Johannink, S. Bahl, A. Nair, J. Luo, A. Kumar, M. Loskyll, J.A. Ojea, E. Solowjow, S. Levine, Residual Reinforcement Learning for Robot Control, 2019 International Conference on Robotics and Automation (ICRA), 6023–6029, 2019.
    [15] L.P. Kaelbling, M.L. Littman, A.W. Moore, Reinforcement learning: A survey, Journal of Artificial Intelligence Research, 4, 237–285, 1996.
    [16] S. Kajita, H. Hirukawa, K. Harada, K. Yokoi, Introduction to Humanoid Robotics, Springer, Berlin, Heidelberg, 2014.
    [17] S. Kajita, F. Kanehiro, K. Kaneko, K. Fujiwara, K. Harada, K. Yokoi, H. Hirukawa, Biped walking pattern generation by using preview control of zero-moment point, 2003 IEEE International Conference on Robotics and Automation (Cat. No.03ch37422), 2, 1620–1626, 2003.
    [18] S. Kajita, M. Morisawa, K. Miura, S. Nakaoka, K. Harada, K. Kaneko, F. Kanehiro, K. Yokoi, Biped walking stabilization based on linear inverted pendulum tracking, Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 4489–4496, 2010.
    [19] R.C. Luo, C.-C. Chen, Biped Walking Trajectory Generator Based on Three-Mass With Angular Momentum Model Using Model Predictive Control, IEEE Transactions on Industrial Electronics, 63, 1, 268–276, 2016.
    [20] K. Lynch, F. Park, Modern Robotics: Mechanics, Planning, and Control, Cambridge University Press, Cambridge, UK, 2017.
    [21] S. Macenski, T. Foote, B. Gerkey, C. Lalancette, W. Woodall, Robot Operating System 2: Design, architecture, and uses in the wild, Science Robotics, 7, 66, eabm6074, 2022.
    [22] V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, G. State, Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning, 2021.
    [23] H.T. Najm, A. Sabah Al-Araji, N.S. Ahmad, Mobile and Humanoid Robots in Healthcare: A Systematic Review of Real-World Applications and Deployment Challenges, IEEE Access, 14, 77677–77700, 2026.
    [24] NVIDIA Corporation, NVIDIA PhysX SDK 5: A Scalable Multi-Platform Physics Solution, https://github.com/NVIDIA-Omniverse/PhysX, 2025.
    [25] NVIDIA Corporation, ROS & ROS2 Bridge - Isaac Sim 4.2.0 Documentation, https://docs.isaacsim.omniverse.nvidia.com/4.2.0/features/external_communication/ext_omni_isaac_ros_bridge.html, 2025.
    [26] NVIDIA Corporation, Tuning Joint Drive Gains - Isaac Sim 4.5.0 Documentation, https://docs.isaacsim.omniverse.nvidia.com/4.5.0/robot_setup/joint_tuning.html, 2025.
    [27] NVIDIA Corporation, NVIDIA Isaac Sim: Robotics Simulation and Synthetic Data Generation, https://developer.nvidia.com/isaac/sim, 2026.
    [28] NVIDIA Corporation, NVIDIA Isaac Lab, https://developer.nvidia.com/isaac/lab, 2026.
    [29] NVIDIA Corporation, Livestream Clients - Isaac Sim 5.0.0 Documentation, https://docs.isaacsim.omniverse.nvidia.com/5.0.0/installation/manual_livestream_clients.html, 2026.
    [30] Open Robotics, ROS 2 Documentation: URDF Tutorials (Humble), https://docs.ros.org/en/humble/Tutorials/Intermediate/URDF/URDF-Main.html, 2026.
    [31] T. Osa, J. Pajarinen, G. Neumann, J.A. Bagnell, P. Abbeel, J. Peters, An Algorithmic Perspective on Imitation Learning, Foundations and Trends in Robotics, 7, 1–2, 1–179, 2018.
    [32] X.B. Peng, Z. Ma, P. Abbeel, S. Levine, A. Kanazawa, AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control, ACM Transactions on Graphics (TOG), 40, 4, 20, 2021.
    [33] M.-C. Popescu, V.E. Balas, L. Perescu-Popescu, N. Mastorakis, Multilayer perceptron and neural networks, WSEAS Transactions on Circuits and Systems, 8, 7, 579–588, 2009.
    [34] M. Posa, C. Cantu, R. Tedrake, A direct method for trajectory optimization of rigid bodies through contact, The International Journal of Robotics Research, 33, 1, 69–81, 2014.
    [35] A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, S. Levine, Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations, 2018.
    [36] N. Rudin, D. Hoeller, P. Reist, M. Hutter, Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning, Proceedings of the 5th Conference on Robot Learning, 164, 91–100, 2022.
    [37] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal Policy Optimization Algorithms, Arxiv Preprint Arxiv:1707.06347, 2017.
    [38] screamlab, fusion2urdf-ros2: Fusion360 to URDF converter for ROS2, https://github.com/screamlab/fusion2urdf-ros2, 2026.
    [39] screamlab, fusion_xacro2urdf2unity, https://github.com/screamlab/fusion_xacro2urdf2unity, 2026.
    [40] J. Siekmann, K. Green, J. Warila, A. Fern, J. Hurst, Blind Bipedal Stair Traversal via Sim-to-Real Reinforcement Learning, Robotics: Science and Systems, 2021.
    [41] J. Tan, T. Zhang, E. Coumans, A. Iscen, Y. Bai, D. Hafner, S. Bohez, V. Vanhoucke, Sim-to-Real: Learning Agile Locomotion For Quadruped Robots, Proceedings of Robotics: Science and Systems, 2018.
    [42] G. Tevatia, S. Schaal, Inverse kinematics for humanoid robots, Proceedings 2000 ICRA. Millennium Conference. IEEE International Conference on Robotics and Automation. Symposia Proceedings (Cat. No.00ch37065), 1, 294–299, 2000.
    [43] X. Wen, L. Wang, Y. Tao, H. Lai, H. Liu, Reinforcement Learning-Based Adaptive Motion Control of Humanoid Robots on Multi-Terrain, Applied Sciences, 16, 5, 2371, 2026.
    [44] P.-B. Wieber, Trajectory Free Linear Model Predictive Control for Stable Walking in the Presence of Strong Perturbations, Proceedings of the 6th IEEE-RAS International Conference on Humanoid Robots (Humanoids), 137–142, 2006.
    [45] K. Yin, K. Loken, M. van de Panne, SIMBICON: Simple Biped Locomotion Control, ACM Transactions on Graphics (TOG), 26, 3, 105, 2007.
    [46] T. Zhang, B. Zheng, R. Nai, Y. Hu, Y.-J. Wang, G. Chen, F. Lin, J. Li, C. Hong, K. Sreenath, others, Hub: Learning extreme humanoid balance, Arxiv Preprint Arxiv:2505.07294, 2025.
    [47] W. Zhao, J.P. Queralta, T. Westerlund, Sim-to-Real Transfer in Deep Reinforcement Learning for Robotics: a Survey, 2020 IEEE Symposium Series on Computational Intelligence (SSCI), 737–744, 2020.

    下載圖示
    校外:立即公開
    QR CODE