| 研究生: |
劉祐任 Liou, You-Ren |
|---|---|
| 論文名稱: |
結合模型預測控制與殘差強化學習之混合控制架構實現穩定雙足機器人步行 A Hybrid Control Architecture for Stable Bipedal Walking Combining MPC and Residual Reinforcement Learning |
| 指導教授: |
鍾俊輝
Chung, Chun-hui |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 機械工程學系 Department of Mechanical Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 90 |
| 中文關鍵詞: | 雙足機器人 、殘差強化學習 、混合控制 |
| 外文關鍵詞: | Bipedal Robot, Residual Reinforcement Learning, Hybrid Control |
| 相關次數: | 點閱:90 下載:1 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
在機器人研究領域中,雙足機器人於複雜動態環境下之穩定行走仍是一項重大挑戰。基於降階動力學模型之傳統控制策略,雖因遵循動力學約束而具備理論穩定性,但在實際應用場景中,其依賴簡化模型之特性容易與真實環境產生未建模誤差,進而導致系統穩定性下降。另一方面,無模型之純強化學習控制策略雖能靈活應對真實環境中之複雜條件,展現較佳的環境適應力,但其產出之控制指令往往缺乏物理基礎,導致此類端到端的策略難以提供理論穩定性保證。為突破上述單一方法之瓶頸,本研究提出一種混合控制架構,旨在完美結合模型控制器之理論穩定性與無模型強化學習之環境適應性,並以參考 Cassie 進行微縮設計之小型雙足機器人為實驗平台,成功於模擬空間中實現穩定且具抗擾能力之行走。在控制架構上,本研究以角動量線性倒立擺模型作為動力學核心,並採用模型預測控制器作為基礎之步態規劃器。透過二次規劃求解,模型預測控制器得以即時計算出符合物理約束之最佳落腳點,作為系統之基準控制輸出。同時,為克服傳統降階線性模型在面臨非線性物理動態時所產生之誤差,本研究進一步引入殘差強化學習框架。在此架構下,殘差強化學習神經網路依據機器人當前之觀測狀態與目標速度誤差,輸出針對模型預測控制器輸出的殘差補償量;系統最終之目標落腳點,即由模型預測控制器輸出之基準輸出與強化學習生成之殘差補償疊加而成。在學習演算法配置上,本研究採用基於演員-評論家架構之近端策略最佳化演算法進行神經網路訓練。為確保控制器具備抗擾能力,神經網路之整體訓練與測試過程皆建構於注入感測器雜訊、訊號延遲以及隨機姿態擾動的模擬環境中進行。於驗證階段,本研究分別於無干擾之理想環境,以及加入上述擾動之干擾環境中,針對混合控制器與模型預測控制器進行速度追蹤測試,藉此比較兩種控制策略之性能;此外,亦探討混合控制器之殘差強化學習模型與純強化學習控制器在訓練效率上之差異。實驗結果顯示,在理想環境中,混合控制器成功消除了模型預測控制器因依賴簡化模型所引發之穩態誤差;在干擾環境中,混合控制器更成功克服了模型預測控制器易失去穩定性而導致狀態發散之致命缺陷。在 X 軸與 Y 軸目標速度指令連續且頻繁切換之動態條件下,混合控制器仍能維持穩健之步態而不失穩,展現其泛化能力。本研究提出之混合控制架構不僅在控制強健性與泛化能力表現優異,其訓練收斂效率與最終策略之存活表現,亦顯著優於純強化學習模型,印證此架構實務部署的潛力。
Achieving stable bipedal locomotion in complex dynamic environments remains a significant challenge. Traditional controllers based on reduced-order models (ROMs) offer theoretical stability but suffer from unmodeled dynamics. Conversely, model-free reinforcement learning (RL) adapts well to complex conditions but lacks physical constraints and stability guarantees. To bridge this gap, this study proposes a hybrid control architecture that integrates the theoretical stability of model-based control with the environmental adaptability of RL, validated in simulation on a miniaturized Cassie-like bipedal robot. The framework utilizes an Angular Momentum Linear Inverted Pendulum Model (ALIPM). A Model Predictive Controller (MPC) serves as the base gait planner, computing physically constrained optimal footholds via quadratic programming. To compensate for the linear ROM's unmodeled nonlinear dynamics, a Residual Reinforcement Learning (RRL) framework is introduced. Trained via Proximal Policy Optimization (PPO), the RRL network calculates residual foothold adjustments based on current observations and velocity tracking errors. The final commanded foothold combines the MPC baseline with the RRL compensation. To ensure robustness, training and evaluation incorporated sensor noise, signal delays, and random perturbations. Results demonstrate that the hybrid controller eliminates the steady-state errors of pure MPC in ideal conditions and effectively prevents state divergence in perturbed environments. It maintains robust gaits even under continuous multidirectional command transitions. Furthermore, the hybrid framework exhibits significantly faster training convergence and higher survival rates compared to pure RL approaches, highlighting its superior robustness, generalization, and substantial potential for real-world deployment.
[1] K. Hirai, M. Hirose, Y. Haikawa, and T. Takenaka, "The development of Honda humanoid robot," in Proceedings. 1998 IEEE international conference on robotics and automation (Cat. No. 98CH36146), 1998, vol. 2: IEEE, pp. 1321–1326.
[2] B.-K. Cho, S.-S. Park, and J.-h. Oh, "Controllers for running in the humanoid robot, HUBO," in 2009 9th IEEE-RAS International Conference on Humanoid Robots, 2009: IEEE, pp. 385–390.
[3] Boston Dynamics. "History." https://bostondynamics.com/about/history/ (accessed July 11, 2026).
[4] Oregon State University. "Bipedal robot developed at Oregon State achieves Guinness World Record in 100 meters." https://news.oregonstate.edu/news/bipedal-robot-developed-oregon-state-achieves-guinness-world-record-100-meters (accessed July 11,, 2026).
[5] Agility Robotics. "Beyond the Hype." https://www.agilityrobotics.com/content/beyond-the-hype (accessed July 11,, 2026).
[6] 本田技研工業株式会社. "2000年代 チャレンジの軌跡." https://global.honda/jp/guide/history-digest/2000/?from=history-digest (accessed July 11, 2026).
[7] S. Kuindersma et al., "Optimization-based locomotion planning, estimation, and control design for the atlas humanoid robot," Autonomous robots, vol. 40, no. 3, pp. 429–455, 2016.
[8] J. Hwangbo et al., "Learning agile and dynamic motor skills for legged robots," Science robotics, vol. 4, no. 26, p. eaau5872, 2019.
[9] M. Vukobratović and D. Juričić, "Contribution to the synthesis of biped gait," IFAC Proceedings Volumes, vol. 2, no. 4, pp. 469–478, 1968.
[10] M. Vukobratović and B. Borovac, "Zero-moment point—thirty five years of its life," International journal of humanoid robotics, vol. 1, no. 01, pp. 157–173, 2004.
[11] S. Kajita, F. Kanehiro, K. Kaneko, K. Yokoi, and H. Hirukawa, "The 3D linear inverted pendulum mode: A simple modeling for a biped walking pattern generation," in Proceedings 2001 IEEE/RSJ International Conference on Intelligent Robots and Systems. Expanding the Societal Role of Robotics in the the Next Millennium (Cat. No. 01CH37180), 2001, vol. 1: IEEE, pp. 239–246.
[12] Y. Gong and J. Grizzle, "One-step ahead prediction of angular momentum about the contact point for control of bipedal locomotion: Validation in a lip-inspired controller," in 2021 IEEE International Conference on Robotics and Automation (ICRA), 2021: IEEE, pp. 2832–2838.
[13] J. Pratt, J. Carff, S. Drakunov, and A. Goswami, "Capture point: A step toward humanoid push recovery," in 2006 6th IEEE-RAS international conference on humanoid robots, 2006: Ieee, pp. 200–207.
[14] J. Englsberger, C. Ott, and A. Albu-Schäffer, "Three-dimensional bipedal walking control based on divergent component of motion," Ieee transactions on robotics, vol. 31, no. 2, pp. 355–368, 2015.
[15] P.-B. Wieber, "Trajectory free linear model predictive control for stable walking in the presence of strong perturbations," in 2006 6th IEEE-RAS International Conference on Humanoid Robots, 2006: IEEE, pp. 137–142.
[16] A. Herdt, H. Diedam, P.-B. Wieber, D. Dimitrov, K. Mombaur, and M. Diehl, "Online walking motion generation with automatic footstep placement," Advanced Robotics, vol. 24, no. 5-6, pp. 719–737, 2010.
[17] G. Gibson, O. Dosunmu-Ogunbi, Y. Gong, and J. Grizzle, "Terrain-adaptive, alip-based bipedal locomotion controller via model predictive control and virtual constraints," in 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2022: IEEE, pp. 6724–6731.
[18] T. McGeer, "Passive dynamic walking," Int. J. Robotics Res., vol. 9, no. 2, pp. 62–82, 1990.
[19] E. R. Westervelt, J. W. Grizzle, and D. E. Koditschek, "Hybrid zero dynamics of planar biped walkers," IEEE transactions on automatic control, vol. 48, no. 1, pp. 42–56, 2003.
[20] J. Siekmann, Y. Godse, A. Fern, and J. Hurst, "Sim-to-real learning of all common bipedal gaits via periodic reward composition," in 2021 IEEE international conference on robotics and automation (ICRA), 2021: IEEE, pp. 7309–7315.
[21] Z. Xie, P. Clary, J. Dao, P. Morais, J. Hurst, and M. Panne, "Learning locomotion skills for cassie: Iterative design and sim-to-real," in Conference on Robot Learning, 2020: PMLR, pp. 317–329.
[22] Z. Li et al., "Reinforcement learning for robust parameterized locomotion control of bipedal robots," in 2021 IEEE International Conference on Robotics and Automation (ICRA), 2021: IEEE, pp. 2811–2817.
[23] X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, "Sim-to-real transfer of robotic control with dynamics randomization," in 2018 IEEE international conference on robotics and automation (ICRA), 2018: IEEE, pp. 3803–3810.
[24] J. Dao, K. Green, H. Duan, A. Fern, and J. Hurst, "Sim-to-real learning for bipedal locomotion under unsensed dynamic loads," in 2022 International Conference on Robotics and Automation (ICRA), 2022: IEEE, pp. 10449–10455.
[25] J. Siekmann, K. Green, J. Warila, A. Fern, and J. Hurst, "Blind bipedal stair traversal via sim-to-real reinforcement learning," arXiv preprint arXiv:2105.08328, 2021.
[26] W. Yu, G. Turk, and C. K. Liu, "Learning symmetric and low-energy locomotion," ACM Transactions on Graphics (TOG), vol. 37, no. 4, pp. 1–12, 2018.
[27] Z. Li, X. B. Peng, P. Abbeel, S. Levine, G. Berseth, and K. Sreenath, "Robust and versatile bipedal jumping control through reinforcement learning," arXiv preprint arXiv:2302.09450, 2023.
[28] H. Duan, J. Dao, K. Green, T. Apgar, A. Fern, and J. Hurst, "Learning task space actions for bipedal locomotion," in 2021 IEEE International Conference on Robotics and Automation (ICRA), 2021: IEEE, pp. 1276–1282.
[29] K. Green, Y. Godse, J. Dao, R. L. Hatton, A. Fern, and J. Hurst, "Learning spring mass locomotion: Guiding policies with a reduced-order model," IEEE Robotics and Automation Letters, vol. 6, no. 2, pp. 3926–3932, 2021.
[30] R. Cheng, G. Orosz, R. M. Murray, and J. W. Burdick, "End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks," in Proceedings of the AAAI conference on artificial intelligence, 2019, vol. 33, no. 01, pp. 3387–3395.
[31] T. Johannink et al., "Residual reinforcement learning for robot control," in 2019 international conference on robotics and automation (ICRA), 2019: IEEE, pp. 6023–6029.
[32] S. H. Bang, C. A. Jové, and L. Sentis, "Rl-augmented mpc framework for agile and robust bipedal footstep locomotion planning and control," in 2024 IEEE-RAS 23rd International Conference on Humanoid Robots (Humanoids), 2024: IEEE, pp. 607–614.
[33] G. A. Castillo, B. Weng, W. Zhang, and A. Hereid, "Robust feedback motion policy design using reinforcement learning on a 3d digit bipedal robot," in 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021: IEEE, pp. 5136–5143.
[34] G. A. Castillo, B. Weng, S. Yang, W. Zhang, and A. Hereid, "Template model inspired task space learning for robust bipedal locomotion," in 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2023: IEEE, pp. 8582–8589.
[35] S. Kajita et al., "Biped walking pattern generation by using preview control of zero-moment point," in 2003 IEEE international conference on robotics and automation (Cat. No. 03CH37422), 2003, vol. 2: IEEE, pp. 1620–1626.
[36] E. F. Camacho, D. R. Ramírez, D. Limón, D. M. De La Peña, and T. Alamo, "Model predictive control techniques for hybrid systems," Annual reviews in control, vol. 34, no. 1, pp. 21–31, 2010.
[37] K. J. Åström and B. Wittenmark, Computer-controlled systems: theory and design. Courier Corporation, 2013.
[38] D. Q. Mayne, J. B. Rawlings, C. V. Rao, and P. O. Scokaert, "Constrained model predictive control: Stability and optimality," Automatica, vol. 36, no. 6, pp. 789–814, 2000.
[39] R. E. Kalman, "Contributions to the theory of optimal control," Bol. soc. mat. mexicana, vol. 5, no. 2, pp. 102–119, 1960.
[40] R. Bellman, "The theory of dynamic programming," Bulletin of the American Mathematical Society, vol. 60, no. 6, pp. 503–515, 1954.
[41] P. Lancaster and L. Rodman, Algebraic riccati equations. Clarendon press, 1995.
[42] T. Takenaka, T. Matsumoto, and T. Yoshiike, "Real time motion generation and control for biped robot-1 st report: Walking gait pattern generation," in 2009 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2009: IEEE, pp. 1084–1091.
[43] M. B. Popovic, A. Goswami, and H. Herr, "Ground reference points in legged locomotion: Definitions, biological trajectories and control implications," The international journal of robotics research, vol. 24, no. 12, pp. 1013–1032, 2005.
[44] T. Koolen, T. De Boer, J. Rebula, A. Goswami, and J. Pratt, "Capturability-based analysis and control of legged locomotion, part 1: Theory and application to three simple gait models," The international journal of robotics research, vol. 31, no. 9, pp. 1094–1113, 2012.
[45] B. J. Stephens and C. G. Atkeson, "Push recovery by stepping for humanoid robots with force controlled joints," in 2010 10th IEEE-RAS International conference on humanoid robots, 2010: IEEE, pp. 52–59.
[46] E. R. Westervelt, J. W. Grizzle, C. Chevallereau, J. H. Choi, and B. Morris, Feedback control of dynamic bipedal robot locomotion. CRC press, 2018.
[47] T. Apgar, P. Clary, K. Green, A. Fern, and J. W. Hurst, "Fast Online Trajectory Optimization for the Bipedal Robot Cassie," in Robotics: Science and Systems, 2018, vol. 101: Pittsburgh, Pennsylvania, USA, p. 14.
[48] L. Biagiotti and C. Melchiorri, "Trajectory planning," in Trajectory Planning for Automatic Machines and Robots: Springer, 2008, pp. 1–12.
[49] G. E. Farin, Curves and surfaces for CAGD: a practical guide. Morgan Kaufmann, 2002.
[50] J. Nocedal and S. J. Wright, Numerical optimization. Springer, 2006.
[51] K. Levenberg, "A method for the solution of certain non-linear problems in least squares," Quarterly of applied mathematics, vol. 2, no. 2, pp. 164–168, 1944.
[52] D. W. Marquardt, "An algorithm for least-squares estimation of nonlinear parameters," Journal of the society for Industrial and Applied Mathematics, vol. 11, no. 2, pp. 431–441, 1963.
[53] C. W. Wampler, "Manipulator inverse kinematic solutions based on vector formulations and damped least-squares methods," IEEE Transactions on Systems, Man, and Cybernetics, vol. 16, no. 1, pp. 93–101, 1986.
[54] A. Liegeois, "Automatic supervisory control of the configuration and behavior of multibody mechanisms," IEEE transactions on systems, man, and cybernetics, vol. 7, no. 12, pp. 868–871, 1977.
[55] S. B. Slotine and B. Siciliano, "A general framework for managing multiple tasks in highly redundant robotic systems," in proceeding of 5th International Conference on Advanced Robotics, 1991, vol. 2, pp. 1211–1216.
[56] R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction (no. 1). MIT press Cambridge, 1998.
[57] J. Kober, J. A. Bagnell, and J. Peters, "Reinforcement learning in robotics: A survey," The International Journal of Robotics Research, vol. 32, no. 11, pp. 1238–1274, 2013.
[58] J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, "Trust region policy optimization," in International conference on machine learning, 2015: PMLR, pp. 1889–1897.
[59] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, "Proximal policy optimization algorithms," arXiv preprint arXiv:1707.06347, 2017.
[60] J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, "High-dimensional continuous control using generalized advantage estimation," arXiv preprint arXiv:1506.02438, 2015.
[61] V. Konda and J. Tsitsiklis, "Actor-critic algorithms," Advances in neural information processing systems, vol. 12, 1999.
[62] R. S. Sutton, "Learning to predict by the methods of temporal differences," Machine learning, vol. 3, no. 1, pp. 9–44, 1988.
[63] T. Silver, K. Allen, J. Tenenbaum, and L. Kaelbling, "Residual policy learning," arXiv preprint arXiv:1812.06298, 2018.
[64] D. Ho. "Legolas - an open source biped." GitHub. https://github.com/daviddoo02/Legolas-an-open-source-biped (accessed July 21, 2026).