| 研究生: |
郭晏竹 Kuo, Yen-Chu |
|---|---|
| 論文名稱: |
應用動態運動原語於軌跡生成之基於視覺之機械手臂模仿學習研究 Study on Vision-Based Imitation Learning for Robotic Manipulators Using Dynamic Movement Primitives for Trajectory Generation |
| 指導教授: |
鄭銘揚
Cheng, Ming-Yang |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 146 |
| 中文關鍵詞: | 桌邊服務 、機械手臂 、動態運動原語 、視覺模仿學習 、CNN-LSTM |
| 外文關鍵詞: | Table-Side Service, Robotic Manipulator, Dynamic Movement Primitives, Vision-Based Imitation Learning, CNN-LSTM |
| 相關次數: | 點閱:130 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
人工智慧技術與機器人產業結合使具自主感知與決策能力之服務型機器人逐漸導入日常生活,其中餐廳服務為重要應用之一。基於此背景,本論文提出一套機器人桌邊服務系統,利用動態運動原語(DMP)進行軌跡規劃與生成,並結合基於視覺之模仿學習方法,使機械手臂模仿人類視覺感知與動作間的映射關係,進而自主執行桌邊服務任務。本論文之機器人桌邊服務系統分為桌面物件收拾與桌面髒污擦拭兩項任務。在髒污擦拭方面,考量人類擦拭軌跡為往復之週期性運動,採用節律型 DMP 與 MDMP 模仿擦拭動作,並根據髒污辨識結果即時調整擦拭振幅與移動距離,使機械手臂針對不同髒污位置與範圍完成擦拭。在物件收拾方面,為解決模仿學習對高品質訓練資料之需求,本論文提出基於四元數之方向型 DMP 生成訓練軌跡,有效降低人工蒐集資料造成之雜訊影響。接著,本論文訓練結合 YOLOv9 目標決策模組之 CNN-LSTM 夾取策略模型,使系統能完成目標辨識與夾取軌跡預測,並達約八成之成功率。最後,整合兩項任務為完整機器人桌邊服務系統,驗證其可有效完成桌邊服務流程。
With the rapid development of artificial intelligence technologies and the robotics industry, service robots with autonomous perception and decision-making capabilities have gradually been introduced into daily living environments. One of these important applications is restaurant service. Based on this background, this thesis proposes a tableside service system that utilizes Dynamic Movement Primitives (DMPs) for trajectory planning and generation. By incorporating a vision-based imitation learning approach, the robotic arm can mimic the mapping relationship between human visual perception and actions, thereby autonomously executing table-side service tasks. The proposed system mainly consists of two tasks: table clearing and tabletop stain wiping. For the stain wiping task, considering that human wiping trajectories are reciprocal and periodic motions, rhythmic DMPs and MDMP are adopted to imitate wiping movements. Based on the stain recognition results, the wiping amplitude and moving distance are adjusted in real time, allowing the robotic arm to complete wiping motions according to different stain locations and areas. For table clearing, to address the high demand for high-quality training data in imitation learning, this thesis proposes using quaternion-based orientation DMPs to generate training trajectories, which effectively mitigates the impact of noise caused by manual data collection. Subsequently, a CNN-LSTM grasping strategy model integrated with a YOLOv9 target decision-making module is trained, enabling the system to achieve target recognition and grasping trajectory prediction with a success rate of approximately 80%. Finally, the two tasks are integrated into a comprehensive table-side service system, verifying its effectiveness in completing the overall table-side service workflow.
[1] 2025/2026 產業技術白皮書,經濟部技術處,2025。
[2] International Federation of Robotics reports, “Top 5 Global Robotics Trends 2026,” Jan. 2026. [Online] Available:https://ifr.org/ifr-press-releases/news/top5-global-robotics-trends-2026#downloads
[3] 工業技術研究院新聞中心。經濟部正式成立智慧機器人研發中心,工研院聚焦醫療照護、物流倉儲、餐飲服務及巡檢救災,壯大機器人核心競爭力。檢 索 日 期 2026 年 5 月 19 號 , 取 自 :https://www.itri.org.tw/ListStyle.aspx?DisplayStyle=01_content&SiteID=1&MmmID=1036276263153520257&MGID=115052014114581324
[4] X. Liu, C. Huang, J. Li, W. Wan, and C. Yang, "Two-Stage Grasp Detection Method for Robotics Using Point Clouds and Deep Hierarchical Feature Learning Network," IEEE Transactions on Cognitive and Developmental Systems, vol. 16, no. 2, pp. 720-731, April 2024.
[5] 陳慧慈,基於單一視角點雲之機械手臂物件抓取姿態生成演算法研究,碩士論文,國立成功大學電機工程學系,中華民國,2024。
[6] 張中昀,基於視覺之深度增強式學習應用於機械手臂物件吸取任務研究,碩士論文,國立成功大學電機工程學系,中華民國,2024。
[7] S. Song, A. Zeng, J. Lee, and T. Funkhouser, "Grasping in the Wild: Learning 6DoF Closed-Loop Grasping From Low-Cost Demonstrations," IEEE Robotics and Automation Letters, vol. 5, no. 3, pp. 4978-4985, July 2020.
[8] C. Chen, Y. Liu, Z. Jiang, and Y. He, "Robot Autonomous Grasping and Assembly Skill Learning Based on Deep Reinforcement Learning," The International Journal of Advanced Manufacturing Technology, vol. 130, no. 11, pp. 5233-5249, February 2024.
[9] P. Xie, S. Zhang, M. Cui, S. Hou, S. Zhang, and X. Song, "GAP-RL: Grasps as Points for RL Towards Dynamic Object Grasping," IEEE Robotics and Automation Letters, vol. 10, no. 1, pp. 40-47, January 2025.
[10] P. Chen and W. Lu, "Deep Reinforcement Learning Based Moving Object Grasping," Information Sciences, vol. 565, pp. 62-76, July 2021.
[11] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, "Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor," in Proceedings of the 35th International Conference on Machine Learning, vol. 80, pp. 1861-1870, 2018.
[12] Y. Huang, S. Chen, J. Ye, Y. Sun, and J. Gao, "A Novel Robotic Grasping Method for Moving Objects Based on Multi-Agent Deep Reinforcement Learning," Robotics and Computer-Integrated Manufacturing, vol. 86, p. 102644, April 2024.
[13] L. Röstel, J. S. Ha, D. Tanneberg, and J. Peters, "Composing Dextrous Grasping and In-Hand Manipulation via Scoring with a Reinforcement Learning Critic," in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), pp. 11683-11690, 2025.
[14] M. Zare, P. M. Kebria, A. Khosravi, and S. Nahavandi, "A Survey of Imitation Learning: Algorithms, Recent Developments, and Challenges," IEEE Transactions on Cybernetics, vol. 54, no. 12, pp. 7173-7186, December 2024.
[15] S. An, Z. Meng, C. Tang, Y. Zhou, T. Liu, F. Ding, S. Zhang, Y. Mu, R. Song, W. Zhang, Z.-G. Hou, and H. Zhang, "Dexterous Manipulation Through Imitation Learning: A Survey," IEEE Transactions on Automation Science and Engineering, vol. 23, pp. 1760-1792, 2026.
[16] M. Seo, S. Han, K. Sim, S. H. Bang, C. Gonzalez, L. Sentis, and Y. Zhu, "Deep Imitation Learning for Humanoid Loco-manipulation Through Human Teleoperation," in Proceedings of the IEEE-RAS 22nd International Conference on Humanoid Robots (Humanoids), Austin, TX, USA, pp. 1-8, 2023.
[17] M. Seo, H. A. Park, S. Yuan, Y. Zhu, and L. Sentis, "LEGATO: CrossEmbodiment Imitation Using a Grasping Tool," IEEE Robotics and Automation Letters, vol. 10, no. 3, pp. 2854-2861, March 2025.
[18] M. J. Kim, J. Wu, and C. Finn, "Giving Robots a Hand: Learning Generalizable Manipulation with Eye-in-Hand Human Video Demonstrations," arXiv:2307.05959, 2023.
[19] T. Gao, S. Nasiriany, H. Liu, Q. Yang, and Y. Zhu, "PRIME: Scaffolding Manipulation Tasks With Behavior Primitives for Data-Efficient Imitation Learning," IEEE Robotics and Automation Letters, vol. 9, no. 10, pp. 8322-8329, October 2024.
[20] S. Kareer, D. Patel, R. Punamiya, P. Mathur, S. Cheng, C. Wang, J. Hoffman, and D. Xu, "EgoMimic: Scaling Imitation Learning via Egocentric Video," in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Atlanta, GA, USA, 2025.
[21] T. Z. Zhao, V. Kumar, S. Levine, and C. Finn, "Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware," arXiv:2304.13705, 2023.
[22]J. Orbik, A. Agostini, and D. Lee, "Inverse Reinforcement Learning for Dexterous Hand Manipulation," in Proceedings of the IEEE International Conference on Development and Learning (ICDL), Beijing, China, pp. 1-7, 2021.
[23] A. Deka, C. Liu, and K. P. Sycara, "ARC - Actor Residual Critic for Adversarial Imitation Learning," in Proceedings of the 6th Conference on Robot Learning (CoRL), vol. 205, pp. 1446-1456, 2023.
[24] M. Tavassoli, S. Katyara, M. Pozzi, N. Deshpande, D. G. Caldwell, and D. Prattichizzo, "Learning Skills From Demonstrations: A Trend From Motion Primitives to Experience Abstraction," IEEE Transactions on Cognitive and Developmental Systems, vol. 16, no. 1, pp. 57-74, February 2024.
[25] X. Yao, Y. Chen and B. Tripp, "Improved Generalization of Probabilistic Movement Primitives for Manipulation Trajectories," IEEE Robotics and Automation Letters, vol. 9, no. 1, pp. 287-294, January 2024.
[26] A. Prados, S. Garrido, and R. Barber, "Learning and generalization of taskparameterized skills through few human demonstrations," Engineering Applications of Artificial Intelligence, vol. 133, p. 108310, July 2024.
[27] A. C. Dometios, Y. Zhou, X. S. Papageorgiou, C. S. Tzafestas, and T. Asfour, "Vision-Based Online Adaptation of Motion Primitives to Dynamic Surfaces: Application to an Interactive Robotic Wiping Task," IEEE Robotics and Automation Letters, vol. 3, no. 3, pp. 1410-1417, July 2018.
[28] Y. Zhang, C. Zeng, J. Zhang, and C. Yang, "Wavelet Movement Primitives: A Unified Framework for Learning Discrete and Rhythmic Movements," IEEE Robotics and Automation Letters, vol. 10, no. 4, pp. 3142-3149, April 2025.
[29] A.C. Reddy, “Difference between Denavit-Hartenberg (D-H) Classical and Modified Conventions for Forward Kinematics of Robots with Case Study,” in Proceedings of the International Conference on Advanced Materials and manufacturing Technologies, pp. 267–286, December 2014.
[30] P. M. Kebria, S. Al-Wais, H. Abdi, and S. Nahavandi, ‘‘Kinematic and dynamic modelling of UR5 manipulator,’’ in Proceedings of the IEEE International Conference on Systems, Man, and Cybernetics, pp. 4229–4234, October 2016.
[31] Universal Robots, “DH Parameters for calculations of kinematics and dynamics,” Universal Robots, 2018. [Online] Available: https://www.universalrobots.com/articles/ur/application-installation/dh-parameters-for-calculationsof-kinematics-and-dynamics/
[32] R. Keating and N. J. Cowan, "UR5 Inverse Kinematics," Johns Hopkins University, 2016. [Online] Available: https://tianyusongcom.wordpress.com/wpcontent/uploads/2017/12/ur5_inverse_kinematics.pdf
[33] M. Ben-Ari, “A tutorial on Euler angles and quaternions,” Weizmann Institute of Science, Israel, 2014.
[34]J. Diebel, “Representing attitude: Euler angles, unit quaternions, and rotation vectors,” Matrix, vol. 58, no. 15–16, pp. 1–35, 2006.
[35]J. S. Dai, “Euler–Rodrigues formula variations, quaternion conjugation and intrinsic connections,” Mechanism and Machine Theory, vol. 92, pp. 144–152, May 2015.
[36]J. Ijspeert, J. Nakanishi, H. Hoffmann, P. Pastor, and S. Schaal, “Dynamical movement primitives: Learning attractor models for motor behaviors,” Neural Computation, vol. 25, no. 2, pp. 328–373, February 2013.
[37] M. Ginesi, N. Sansonetto, and P. Fiorini, “Overcoming some drawbacks of dynamic movement primitives,” Robotics and Autonomous Systems, vol. 144, Art. p. 103844, July 2021.
[38] A. Ude, B. Nemec, T. Petrič, and J. Morimoto, “Orientation in Cartesian space dynamic movement primitives,” in Proceedings of the IEEE International Conference on Robotics and Automation, Hong Kong, China, pp. 2997–3004, 2014.
[39] M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proceedings of the 36th International Conference on Machine Learning, Proceedings of Machine Learning Research, vol. 97, pp. 6105–6114, 2019.
[40] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, pp. 770–778, 2016.
[41] C.-Y. Wang, I.-H. Yeh, and H.-Y. M. Liao, “YOLOv9: Learning what you want to learn using programmable gradient information,” in Proceedings of the European Conference on Computer Vision, Cham, Switzerland: Springer Nature Switzerland, pp. 1-21, October 2024.
[42] C.-Y. Wang, H.-Y. M. Liao, and I.-H. Yeh, “Designing network design strategies through gradient path analysis,” arXiv:2211.04800, 2022.
[43] C.-Y. Wang, H.-Y. M. Liao, Y.-H. Wu, P.-Y. Chen, J.-W. Hsieh, and I.-H. Yeh, “CSPNet: A new backbone that can enhance learning capability of CNN,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 390–391, 2020.
[44] D. Chen, D. Kong, J. Li, S. Wang, and B. Yin, “A survey of visual affordance recognition based on deep learning,” IEEE Transactions on Big Data, vol. 9, no. 6, pp. 1458–1476, December 2023.
[45] A. Nguyen, D. Kanoulas, D. G. Caldwell, and N. G. Tsagarakis, “Object-based affordances detection with convolutional neural networks and dense conditional random fields,” in Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, Vancouver, BC, Canada, pp. 5908–5915, 2017.
[46] L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking atrous convolution for semantic image segmentation,” arXiv:1706.05587, 2017.
[47] Y. Song, P. Sun, P. Jin, Y. Ren, Y. Zheng, Z. Li, X. Chu, Y. Zhang, T. Li, and J. Gu, "Learning 6-DoF Fine-Grained Grasp Detection Based on Part Affordance Grounding," IEEE Transactions on Automation Science and Engineering, vol. 22, pp. 15200-15214, May 2025.
[48] F. Yang, D. Luo, W. Chen, J. Lin, J. Cai, K. Yang, Z. Li, and Y. Wang, "MultiKeypoint Affordance Representation for Functional Dexterous Grasping," IEEE Robotics and Automation Letters, vol. 10, no. 10, pp. 10306-10313, October 2025.
[49]J. Sun, A. Curtis, Y. You, Y. Xu, M. Koehle, Q. Chen, S. Huang, L. Guibas, S. Chitta, M. Schwager, and H. Li, "ARCH: Hierarchical Hybrid Learning for Long-Horizon Contact-Rich Robotic Assembly," arXiv:2409.16451, 2025.