簡易檢索 / 詳目顯示

研究生: 柯惟中
Ko, Wei-Chung
論文名稱: 基於柔性策略評價之四足機器人步態學習與實驗驗證
Soft Actor-Critic Based Gaits Learning and Experimental Verification for Quadruped Robots
指導教授: 田思齊
Tien, Szu-Chi
學位類別: 碩士
Master
系所名稱: 工學院 - 機械工程學系
Department of Mechanical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 102
中文關鍵詞: MuJoCoSAC 強化學習四足機器人
外文關鍵詞: Soft Actor-Critic, MuJoCo, quadruped robot
相關次數: 點閱:19下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究旨在學習四足機器人於未知環境中的步態行為,利用 MuJoCo 物理模擬平台與 Stable-Baselines3 的 SAC 強化學習方法,設計獎勵函數以優化四足機器人的行走控制策略。為克服模擬與現實間的落差,本研究採用系統識別技術,將實機硬體特性還原至物理模擬引擎中以建立高擬真模型。
    研究結果顯示,SAC 演算法能有效在高維度狀態空間中,為複雜動力系統產生強健的控制策略。此外,將模擬訓練出的最優策略實際部署於真實的四足機器人上,實機已能成功展現出步態前進行為。儘管實機測試中部分感測數據受現實環境干擾而存在誤差,但整體的動態表現仍成功驗證了結合系統識別與 SAC 演算法的控制策略應用至真實環境之可行性。

    This study proposes a reinforcement learning-based control framework for quadruped robots, integrating the Soft Actor-Critic (SAC) algorithm with the MuJoCo physics engine to achieve stable walking gaits and bridge the Sim-to-Real gap. Traditional controllers struggle with nonlinear dynamics and environmental uncertainties, whereas SAC employs a maximum entropy framework to balance exploration and exploitation, effectively preventing premature convergence in high-dimensional state spaces. To minimize the dynamic discrepancy between simulation and the physical world, system identification is performed on the P1S servo motors of the Petoi Bittle quadruped robot, achieving a 93.1$%$ fit for the dynamic response. A target trajectory generator based on Fourier series and a comprehensively designed reward function are introduced to guide the policy network. The simulated results demonstrate that the robot successfully learns a stable diagonal gait, maintaining consistent trunk height and accurately tracking the target linear velocity. Real-world experiments using a distributed architecture (Raspberry Pi and NyBoard) validate the feasibility of the proposed method, showing that the trained policy can be deployed to the physical robot to achieve balanced walking without falling. Overall, this research provides a robust Sim-to-Real transfer methodology for quadruped locomotion.

    摘要 i Extend Abstract ii 致謝 x 圖目錄 xiii 表目錄 xv 符號表 xvi 第一章 緒論 1 第二章 SAC 強化學習演算法 5 2.1 SAC 核心概念 5 2.2 與監督/非監督學習的差異及其更新原理 11 第三章 模擬環境設定 20 3.1 模擬實機架構 21 3.2 整合訓練環境設定 28 第四章 實驗平台與實驗方法 40 4.1 實驗用四足機器人介紹 41 4.2 電路系統與數據傳輸 43 4.3 馬達特性分析 48 4.4 樹莓派系統整合 52 第五章 實驗與討論 57 5.1 模擬結果 57 5.2 模擬狀態輸入之實機測試結果 66 5.3 邊緣運算平台之硬體資源與算力限制 72 第六章 結論與未來展望 73 6.1 結論 73 6.2 未來展望 73 參考文獻 74

    [1] A. Agha et al. Nebula: Quest for robotic autonomy in challenging environments; team costar at the darpa subterranean challenge. arXiv preprint arXiv:2103.11470, 2021.
    [2] L. Pallottino. Robotics for warehouses and logistics: Technologies, challenges, and future directions. Annual Review of Control, Robotics, and Autonomous Systems, 9, 2005.
    [3] M. Milburn et al. Computer-vision based real time waypoint generation for autonomous vineyard navigation with quadruped robots. In 2023 IEEE International Conference on Autonomous Robot Systems and Competitions (ICARSC), pages 239-244. IEEE, 2023.
    [4] K. Yano et al. Development of helios ix: an arm-equipped tracked vehicle. Journal of Robotics and Mechatronics, 23(6):1031-1040, 2011.
    [5] 李彥廷. 雙點主動地形感知之四足機器人強化學習運動控制. 碩士論文, 國立臺灣科技大學自動化及控制研究所, 2025.
    [6] 鄭子嘉. 基於模型預測控制與機器學習參數優化之四足機器狗模擬. 國立臺灣大學機械工程學系學位論文, pages 1-105, 2025.
    [7] H. Li et al. Quadruped robots: Briefing mechanical design, control, and applications. Robotics, 14(5):57, 2025.
    [8] A. P. Miller, Fangzhou Yu, et al. High-performance reinforcement learning on spot: Optimizing simulation parameters with distributional measures. Preprint or Conference Proceedings, 2025. Research on Sim-to-Real RL deployment on Boston Dynamics Spot.
    [9] A. M. El-Dalatony et al. Cascaded pid trajectory tracking control for quadruped robotic leg. International Journal of Mechanical Engineering and Robotics Research, 12(1):40-47, 2023.
    [10] J. Di Carlo et al. Dynamic locomotion in the mit cheetah 3 through convex model-predictive control. In 2018 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 1-9. IEEE, 2018.
    [11] D. C. Tosun et al. Comparison of pid and lqr controllers on a quadrotor helicopter. International journal of systems applications, engineering & development, 9:136-143, 2015.
    [12] M. A. Wiering and M. Van Otterlo. Reinforcement learning. Adaptation, learning, and optimization, 12(3):729, 2012.
    [13] J. Schulman et al. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
    [14] G. Barth-Maron et al. Distributed distributional deterministic policy gradients. arXiv preprint arXiv:1804.08617, 2018.
    [15] T. Haarnoja et al. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning, pages 1861-1870. PMLR, 2018.
    [16] L. Angel et al. Adams/matlab co-simulation: Dynamic systems analysis and control tool. Applied Mechanics and Materials, 232:527-531, 2012.
    [17] S. L. Delp et al. Opensim: open-source software to create and analyze dynamic simulations of movement. IEEE transactions on biomedical engineering, 54(11):1940-1950, 2007.
    [18] E. Todorov et al. Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ international conference on intelligent robots and systems, pages 5026-5033. IEEE, 2012.
    [19] J. Tobin et al. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 23-30. IEEE, 2017.
    [20] A. Raffin et al. Stable-baselines3: Reliable reinforcement learning implementations. Journal of machine learning research, 22(268):1-8, 2021.
    [21] C.-J. Zhang. Using Soft Actor-Critic Learning to Achieve Dynamic Maneuvers of Quadruped Robots. Master's thesis, Department of Mechanical Engineering, National Cheng Kung University, Tainan, Taiwan, 2024.
    [22] V. Nair and G. E. Hinton. Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th international conference on machine learning (ICML-10), pages 807-814, 2010.
    [23] H. Van Hasselt et al. Deep reinforcement learning with double q-learning. In Proceedings of the AAAI conference on artificial intelligence, volume 30, 2016.
    [24] S. Fujimoto et al. Addressing function approximation error in actor-critic methods. In International conference on machine learning, pages 1587-1596. PMLR, 2018.
    [25] B. T. Polyak and A. B. Juditsky. Acceleration of stochastic approximation by averaging. SIAM journal on control and optimization, 30(4):838-855, 1992.
    [26] T. P. Lillicrap et al. Continuous control with deep reinforcement learning, September 15 2020. US Patent 10,776,692.
    [27] J. Schulman et al. Trust region policy optimization. In International conference on machine learning, pages 1889-1897. PMLR, 2015.
    [28] S. J. Pan and Q. Yang. A survey on transfer learning. IEEE Transactions on Knowledge and Data Engineering, 22(10):1345-1359, 2010.
    [29] V. Mnih et al. Human-level control through deep reinforcement learning. nature, 518(7540):529-533, 2015.
    [30] D. P. Kingma and M. Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
    [31] J. Cai et al. Modeling method of autonomous robot manipulator based on d-h algorithm. Mobile Information Systems, 2021(1):4448648, 2021.
    [32] E. Todorov. Convex and analytically-invertible dynamics with contacts and constraints: Theory and implementation in mujoco. In 2014 IEEE International Conference on Robotics and Automation (ICRA), pages 6054-6061. IEEE, 2014.
    [33] MuJoCo Developers. Mujoco documentation: Constraint solver, 2026. Accessed: 2026-05-25.
    [34] E. Todorov. A convex, smooth and invertible contact model for trajectory optimization. In 2011 IEEE International Conference on Robotics and Automation, pages 1071-1076. IEEE, 2011.
    [35] E. Todorov. Implicit nonlinear complementarity: A new approach to contact dynamics. In 2010 IEEE international conference on robotics and automation, pages 2322-2329. IEEE, 2010.
    [36] X. B. Peng et al. Deepmimic: Example-guided deep reinforcement learning of physics-based character skills. ACM Transactions On Graphics (TOG), 37(4):1-14, 2018.
    [37] Petoi. Petoi Bittle: Open source bionic robot dog. [https://bittle.petoi.com/](https://bittle.petoi.com/), 2026. [Online; accessed 3-July-2026].
    [38] Petoi. Nyboard v1 1 & nyboard v1 2 - Petoi documentation. [https://docs.petoi.com/nyboard/nyboard-v1_1-and-nyboard-v1_2](https://docs.petoi.com/nyboard/nyboard-v1_1-and-nyboard-v1_2), 2026. [Online; accessed 3-July-2026].
    [39] InvenSense Inc. Mpu-6000 and mpu-6050 product specification revision 3.4. Datasheet, InvenSense Inc., Sunnyvale, CA, USA, 2013. Accessed: Jul. 7, 2026.
    [40] C. Fuentes-Silva et al. Constrained gray-box identification of electromechanical systems under unfiltered step-response data. Information, 16(12):1079, 2025.
    [41] P. W. Hodges et al. Coexistence of stability and mobility in postural control: evidence from postural compensation for respiration. Experimental brain research, 144(3):293-302, 2002.
    [42] S. O. Madgwick et al. An efficient orientation filter for inertial and inertial/magnetic sensor arrays. 2010.
    [43] Department of Mechanical Engineering, University of Utah. Me en 3200 mechatronics i lab: Motors. [https://my.mech.utah.edu/~me3200/labs/motors.pdf](https://www.google.com/search?q=https://my.mech.utah.edu/~me3200/labs/motors.pdf), 2026. Accessed: 2026-07-21.

    下載圖示
    校外:立即公開
    QR CODE