| 研究生: |
張宸嘉 Zhang, Chen-Jia |
|---|---|
| 論文名稱: |
運用柔性策略-評價學習以實現四足機器人之動態動作 Using Soft Actor-Critic Learning to Achieve Dynamic Maneuvers of Quadruped Robots |
| 指導教授: |
田思齊
Tien, Szu-Chi |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 機械工程學系 Department of Mechanical Engineering |
| 論文出版年: | 2024 |
| 畢業學年度: | 112 |
| 語文別: | 中文 |
| 論文頁數: | 94 |
| 中文關鍵詞: | MuJoCo 、SAC強化學習 、四足機器人 、獎勵機制 |
| 外文關鍵詞: | MuJoCo, SAC reinforcement learning, quadruped robot, reward system |
| 相關次數: | 點閱:124 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究利用MuJoCo模擬平台和Stable Baselines3的SAC強化學習方法,設計合理的獎勵函數優化四足機器人之運動控制策略,使其站立至目標高度。除了對模擬過程中不同的獎勵機制進行了詳細分析,還將最優模擬結果中獲取的機器人關節角度變化應用於真實的Petoi Bittle四足機器人。研究結果顯示,獎勵機制的設計對強化學習效果有著重要影響,直接關係到機器人運動控制的性能和穩定性。此外,比較模擬和實驗中的高度和姿態角度變化,發現兩者之結果相似,可見模擬中所得的控制策略可用於實際環境中。總體而言,使用SAC強化學習的方式能夠在高維度狀態空間下為複雜動力學系統產生足以面對未知環境的控制策略,並實際驗證模擬結果應用至真實四足機器人的可行性。
In this research, an appropriate reward function is designed to optimise the motion control strategy of a quadruped robot to stand up to the target height using the MuJoCo simulation platform and the SAC reinforcement learning method of Stable Baselines3. In addition to the detailed analysis of different reward functions in the simulation process, the variation of robot joint angles obtained from the optimal simulation results is also applied to the real Petoi Bittle quadruped robot. The results show that the design of the reward mechanism has a significant impact on the reinforcement learning effect, which is directly related to the performance and stability of the robot's motion control. Furthermore, when comparing the height and attitude angle variations between the simulation and the experiment, we found that the results are similar, which shows that the control strategies obtained in the simulation can be used in the real environment. Overall, the SAC reinforcement learning approach is able to generate control strategies for complex dynamical systems in a high-dimensional state space that are adequate for unknown environments, and the feasibility of applying the simulation results to real quadruped robots can be practically verified.
[1] Peng Xiao et al. Design of environment perception system for quadruped in spection robot. In 2023 IEEE 11th Joint International Information Technol ogy and Artificial Intelligence Conference (ITAIC), volume 11, pages 598–602. IEEE, 2023.
[2] Xiyun Jin et al. A vision perception module on quadruped robot in metro inspection. In 2022 6th International Conference on Automation, Control and Robots (ICACR), pages 12–16. IEEE, 2022.
[3] Yunjie Zhou et al. Automatic inspection method of cable tunnel in complex environment based on quadruped robot. In 2021 IEEE 3rd International Conference on Frontiers Technology of Information and Computer (ICFTIC), pages 599–603. IEEE, 2021.
[4] Frank E Schneider and Dennis Wildermuth. Assessing the search and rescue domain as an applied and realistic benchmark for robotic systems. In 2016 17th International Carpathian Control Conference (ICCC), pages 657–662. IEEE, 2016.
[5] Markus Eich et al. A versatile stair-climbing robot for search and rescue applications. In 2008 IEEE international workshop on safety, security and rescue robotics, pages 35–40. IEEE, 2008.
[6] Lee Milburn et al. Computer-vision based real time waypoint generation for autonomous vineyard navigation with quadruped robots. In 2023 IEEE International Conference on Autonomous Robot Systems and Competitions (ICARSC), pages 239–244. IEEE, 2023.
[7] Nianfei Du et al. A comfortable interaction strategy for the visually impaired using quadruped guidance robot. In 2023 42nd Chinese Control Conference (CCC), pages 4849–4853. IEEE, 2023.
[8] Nan Wang et al. A analyse on the virtual simulation of vehicle stability. In 2010 International Conference on Measuring Technology and Mechatronics Automation, volume 1, pages 14–17, 2010.
[9] Yang Ying et al. Control strategy research and simulation analysis of electric power steering system for automobile. In 2009 WRI Global Congress on Intelligent Systems, volume 2, pages 228–232. IEEE, 2009.
[10] Qingchao Wei et al. A dynamic simulation model of linear metro system with admas/rail. In 2007 International Conference on Mechatronics and Automa tion, pages 2037–2042, 2007.
[11] Vahid Noei and Heba Lakany. Analysis of movement of an elbow joint with a wearable robotic exoskeleton using opensim software. In 2022 44th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), pages 4342–4345. IEEE, 2022.
[12] Jonathan Camargo et al. Opensim model for biomechanical analysis with the open-source bionic leg. In 2022 International Symposium on Medical Robotics (ISMR), pages 1–6. IEEE, 2022.
[13] Xinyan Jiang and Istv´an B´ ır´ o. Muscle force estimation in running gait analysis after a prolonged running session via opensim. In 2023 4th International Conference on Computer Engineering and Application (ICCEA), pages 498 502. IEEE, 2023.
[14] Emanuel Todorov. Convex and analytically-invertible dynamics with contacts and constraints: Theory and implementation in mujoco. In 2014 IEEE Inter national Conference on Robotics and Automation (ICRA), pages 6054–6061. IEEE, 2014.
[15] Krishnendu Roy et al. Walking of prismatic knee biped robot using reinforce ment learning. In 2023 IEEE 4th Annual Flagship India Council International Subsections Conference (INDISCON), pages 1–6. IEEE, 2023.
[16] Chengye Liao et al. Performance comparison of typical physics engines using robot models with multiple joints. IEEE Robotics and Automation Letters, 2023.
[17] Petrus Sutyasadi and Manukid Parnichun. Trotting control of a quadruped robot using pid-ilc. In IECON 2015-41st Annual Conference of the IEEE Industrial Electronics Society, pages 004400–004405. IEEE, 2015.
[18] Paolo Arena et al. Mpc-based control strategy of a neuro-inspired quadruped robot. In 2021 International Joint Conference on Neural Networks (IJCNN), pages 1–8. IEEE, 2021.
[19] Yongzhe Du et al. Mpc-based tilting and forward motion control of quadruped robots. In 2022 5th International Symposium on Autonomous Systems (ISAS), pages 1–6. IEEE, 2022.
[20] AA M. Zahir et al. Genetic algorithm optimization of pid controller for brushed dc motor. In Intelligent Manufacturing & Mechatronics: Proceedings of Symposium, 29 January 2018, Pekan, Pahang, Malaysia, pages 427–437. Springer, 2018.
[21] Andri Mirzal et al. Pid parameters optimization by using genetic algorithm. arXiv preprint arXiv:1204.0885, 2012.
[22] Jing Tian et al. A mpc and genetic algorithm based approach for multiple uavs cooperative search. In Computational Intelligence and Security: International Conference, CIS 2005, Xi’an, China, December 15-19, 2005, Proceedings Part I, pages 399–404. Springer, 2005.
[23] Yang Sun et al. Trajectory tracking control design for 4ws vehicle based on particle swarm optimization and phase plane analysis. Applied Sciences, 14(9):3664, 2024.
[24] Luobin Cui and Ying Tang. Comparing the effectiveness of ppo and its vari ants in training ai to play game. In 2023 International Conference on Cyber Physical Social Intelligence (ICCSI), pages 521–526. IEEE, 2023.
[25] Yikang Ouyang et al. Reinforcement learning-based path tracking with ap plication of quadruped robot. In 2021 China Automation Congress (CAC), pages 7069–7074. IEEE, 2021.
[26] Timothy P Lillicrap et al. Continuous control with deep reinforcement learn ing. arXiv preprint arXiv:1509.02971, 2015.
[27] Jingliang Duan et al. Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors. IEEE transactions on neural networks and learning systems, 33(11):6584–6598, 2021.
[28] Tuomas Haarnoja et al. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning, pages 1861–1870. PMLR, 2018.
[29] Antonin Raffin et al. Stable-baselines3: Reliable reinforcement learning im plementations. Journal of Machine Learning Research, 22(268):1–8, 2021.
[30] Petoi. https://www.petoi.com/pages/bittle-open-source-bionic-robot-dog. In Bittle Open Source Bionic Robot Dog. Accessed: 2024-07-16.
[31] Petoi. https://docs.petoi.com/nyboard/nyboard-v1 1-and-nyboard-v1 2. In NyBoard V1 1 & NyBoard V1 2. Accessed: 2024-07-16.
[32] Vinod Nair and Geoffrey E Hinton. Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th international conference on machine learning (ICML-10), pages 807–814, 2010.
[33] Boris T Polyak. Some methods of speeding up the convergence of iteration methods. Ussr computational mathematics and mathematical physics, 4(5):1 17, 1964.
[34] Saeed B Niku. Introduction to robotics: analysis, control, applications. John Wiley & Sons, 2020.
[35] Emanuel Todorov et al. Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ international conference on intelligent robots and systems, pages 5026–5033. IEEE, 2012.
[36] David E Stewart and Jeffrey C Trinkle. An implicit time-stepping scheme for rigid body dynamics with inelastic collisions and coulomb friction. Inter national Journal for Numerical Methods in Engineering, 39(15):2673–2691, 1996.
[37] Emanuel Todorov. Implicit nonlinear complementarity: A new approach to contact dynamics. In 2010 IEEE international conference on robotics and automation, pages 2322–2329. IEEE, 2010.
[38] Naum Zuselevich Shor. Minimization methods for non-differentiable func tions, volume 3. Springer Science & Business Media, 2012.
[39] James Diebel et al. Representing attitude: Euler angles, unit quaternions, and rotation vectors. Matrix, 58(15-16):1–35, 2006.
[40] Evan G Hemingway and Oliver M O’Reilly. Perspectives on euler angle sin gularities, gimbal lock, and the orthogonality of applied forces and applied moments. Multibody system dynamics, 44:31–56, 2018.
[41] Ankit Choudhary. https://www.analyticsvidhya.com/blog/2018/09/reinforcement learning-model-based-planning-dynamic-programming/. In Nuts & Bolts of Reinforcement Learning: Model Based Planning using Dynamic Program ming. Accessed: 2024-07-16.