| 研究生: |
謝頤賢 HSIEN, YI-HSIEN |
|---|---|
| 論文名稱: |
基於數位孿生與多視角視覺之無標記物件定位與機械臂靠近任務 Markerless Object Localization and Robotic Arm Approaching Task Based on Digital Twin and Multi-View Vision |
| 指導教授: |
蘇文鈺
Su, Wen-Yu |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 70 |
| 中文關鍵詞: | 機器視覺 、YOLO26 、ROS 2 、空間三維定位 、Isaac Sim 、達明機械臂 、MoveIt 2 |
| 外文關鍵詞: | Machine Vision, YOLO26, ROS 2, 3D Spatial Localization, Isaac Sim, Techman Robotic Manipulator, MoveIt 2 |
| 相關次數: | 點閱:5 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本論文旨在設計與實作一套基於多攝影機視覺的 3D 物件定位與機械臂控制系統,在無位置標記(marker)情況下,解決物件的空間座標定位及機械手臂移動的任務。本系統以 ROS 2 (Robot Operating System 2) 為核心開發框架,結合 YOLO26 物件偵測技術,達成即時且強健的物件辨識。本論文的一大特點在於,執行過程中不需要依賴深度相機與標記物 (如 ArUco Marker),僅透過 RGB 攝影機即可估算物體座標,並使機械臂靠近目標物體。
在視覺定位演算法方面,本論文實作了一套多攝影機光線交會叢集演算法。系統首先從各攝影機影像中擷取目標的二維像素座標,將其反投影為三維空間射線,再透過叢集與最小平方法計算出最佳的交會點,估測物件的世界座標。此外,我們也整合了多目標追蹤模組,進一步提升連續畫面下物件座標的穩定性。
在系統整合方面,本論文建構了基於 Isaac Sim 數位孿生環境的測試平台。先透過在虛擬環境中進行演算法的快速迭代與測試,隨後部署至實體機器人系統中。本系統不僅完成無標記物件環境下的定位任務,更進一步將定位結果傳遞至 MoveIt 2 運動規劃框架,驅動達明機械手臂 (TM5S) 執行靠近物件的任務。實驗結果驗證了此架構的可行性,證明本系統能有效串連空間定位與運動規劃模組,成功實現無標記視覺定位並導引機器手臂靠近之任務。最後,本論文分別針對數位孿生與實體操作環境,設計相應之誤差評估指標進行分析,驗證了本系統之定位精度。實驗結果顯示,在虛擬環境中估算多類別物體座標之平均誤差約為 1.23 公分(第 95 百分位數誤差約為 3.02 公分);而在實體環境中,本論文以代表性物件進行驗證,機械臂端點至物件中心之距離平均誤差約為 1.67 公分(最大誤差在 3.1 公分以內)。此數據顯示本系統在無位置標記情況下,仍具備引導機械臂靠近目標物與估算物體座標之能力。
This thesis aims to design and implement a multi-camera vision-based 3D object localization and robotic manipulator control system. Under markerless conditions, the system addresses 3D object localization and robotic manipulator movement tasks. This system uses ROS 2 (Robot Operating System 2) as the core development framework and integrates the YOLO26 object detection model to achieve real-time and robust object detection. A major feature of this thesis is that the system does not rely on depth cameras or artificial markers, such as ArUco markers, during operation. Instead, it estimates 3D object coordinates using only RGB cameras, enabling the robotic manipulator to approach the target object.
For 3D object localization, this thesis implements a multi-camera ray-intersection clustering algorithm. The system first extracts the 2D pixel coordinates of the target from each camera image and back-projects them into 3D spatial rays. It then calculates the optimal intersection point using clustering and least squares to estimate the object's world coordinates. In addition, we also integrate a multi-object tracking module to further improve the stability of 3D object coordinates across consecutive frames.
For system integration, this thesis develops a test platform built on the Isaac Sim digital twin environment. Algorithms are first rapidly iterated and tested in the virtual environment, and are then deployed to the physical robotic system. This system not only completes 3D object localization tasks in a markerless object environment but also further transmits the 3D object localization results to the MoveIt 2 motion planning framework, driving the Techman robotic manipulator (TM5S) to perform object-approach tasks. Experimental results validate the feasibility of this architecture and demonstrate that the system can effectively integrate 3D object localization and motion planning modules, thereby enabling markerless 3D object localization and robotic manipulator approach tasks. Finally, this thesis designs corresponding error evaluation metrics for both the digital twin and physical operating environments to analyze and verify the system's 3D object localization accuracy. Experimental results show that, in the virtual environment, the average error of multi-class 3D object-coordinate estimation is approximately 1.23 cm, and the 95th percentile error is approximately 3.02 cm. In the physical environment, representative objects were used for validation, yielding a mean distance error of approximately 1.67 cm between the end-effector tip and the object center, with a maximum error below 3.1 cm. These results show that under markerless conditions, this system can guide the robotic manipulator toward target objects and estimate their 3D coordinates.
[1] Jia Chen, Dongli Wu, Peng Song, Fuqin Deng, Ying He, and Shiyan Pang. Multiview triangulation: Systematic comparison and an improved method. IEEE Access, 8:21017–21027, 2020.
[2] NVIDIA. Isaac Sim. Version 5.1.0 [Computer software]. Available at: https://github.com/isaac-sim/IsaacSim/tree/v5.1.0, 2025. Apache-2.0 license.
[3] Ranjan Sapkota, Rahul Harsha Cheppally, Ajay Sharda, and Manoj Karkee. Yolo26: key architectural enhancements and performance benchmarking for realtime object detection. arXiv preprint arXiv:2509.25164, 2025.
[4] Steven Macenski, Tully Foote, Brian Gerkey, Chris Lalancette, and William Woodall. Robot operating system 2: Design, architecture, and uses in the wild. Science robotics, 7(66):eabm6074, 2022.
[5] David Coleman, Ioan Sucan, Sachin Chitta, and Nikolaus Correll. Reducing the barrier to entry of complex robotic software: a moveit! case study. arXiv preprint arXiv:1404.3785, 2014.
[6] Joaqu´ın Alori, Alan Descoins, javier, Facundo Lezama, KotaYuhara, Diego Fern´andez, Agust´ın Castro, fatih, David, Roc´ıo Cruz Linares, Francisco Kurucz, Braulio R´ıos, shafu.eth, Kadir Nar, David Huh, and Moises. tryolabs/norfair:v2.2.0, January 2023.
[7] Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence, 39(6):1137–1149, 2016.
[8] Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 779–788, 2016.
[9] Mupparaju Sohan, Thotakura Sai Ram, and Ch Venkata Rami Reddy. A review on yolov8 and its advancements. In International conference on data intelligence and cognitive informatics, pages 529–545. Springer, 2024.
[10] Ioan A Sucan, Mark Moll, and Lydia E Kavraki. The open motion planning library. IEEE Robotics & Automation Magazine, 19(4):72–82, 2012.
[11] Techman Robot Inc. TM5S Collaborative Robot. https://www.tm-robot.com/zh-hant/product/tm-ai-cobot-s-series/tm5s.
[12] AVerMedia Technologies, Inc. Live Streamer CAM 513 (PW513). https://www.avermedia.com/tw/product-detail/PW513.
[13] TahirNilin. Apple. Sketchfab. https://sketchfab.com/3d-models/apple-643eb66651864bb78871e5c1066b4ef6, 2020. 3D model, licensed under CC BY 4.0.
[14] firehawksoftware. Banana. Sketchfab. https://sketchfab.com/3d-models/banana-ada7c35a1a5742f1b4c528eb3daee35b, 2020. 3D model, licensed under CC BY-NC 4.0.
[15] Batuhan13. Wine Bottle. Sketchfab. https://sketchfab.com/3d-models/wine-bottle-b0ed331c418d4959a3f050ce88f47346, 2019. 3D model, licensed under CC BY 4.0.
[16] Hannah Eliason. Craft Teddy Bear Photogrammetry. Sketchfab. https://sketchfab.com/3d-models/craft-teddy-bear-photogrammetry-cb81496a28ba423eb5e3bd2149143bc8, 2020. 3D model, licensed under CC BY 4.0.
[17] SAM 3D Team, Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, Aohan Lin, Jiawei Liu, Ziqi Ma, Anushka Sagar, Bowen Song, Xiaodong Wang, Jianing Yang, Bowen Zhang, Piotr Doll´ar, Georgia Gkioxari, Matt Feiszli, and Jitendra Malik. Sam 3d: 3dfy anything in images. 2025.
[18] Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, Jie Lei, Tengyu Ma, Baishan Guo, Arpit Kalla, Markus Marks, Joseph Greer, Meng Wang, Peize Sun, Roman R¨adle, Triantafyllos Afouras, Effrosyni Mavroudi, Katherine Xu, Tsung-Han Wu, Yu Zhou, Liliane Momeni, Rishi Hazra, Shuangrui Ding, Sagar Vaze, Francois Porcher, Feng Li, Siyuan Li, Aishwarya Kamath, Ho Kei Cheng, Piotr Doll´ar, Nikhila Ravi, Kate Saenko, Pengchuan Zhang, and Christoph Feichtenhofer. Sam 3: Segment anything with concepts, 2025.
[19] G. Bradski. The OpenCV Library. Dr. Dobb’s Journal of Software Tools, 2000.
[20] Zhengyou Zhang. A flexible new technique for camera calibration. IEEE Transactions on pattern analysis and machine intelligence, 22(11):1330–1334, 2000.
[21] Techman Robot Inc. TM2 ROS 2 Driver. https://github.com/TechmanRobotInc/tm2_ros2/tree/humble. GitHub repository, humble branch.
[22] Robotiq Inc. IO Coupling TM Model. Robotiq Support download. STEP model.