簡易檢索 / 詳目顯示

研究生: 林聿翔
Lin, Yu-Siang
論文名稱: 結合視觸覺之食物物件自適應智慧夾取系統研究
Research on an Adaptive and Intelligent Grasping System for Food Objects Based on Vision-Tactile Fusion
指導教授: 鍾俊輝
Chung, Chun-hui
學位類別: 碩士
Master
系所名稱: 工學院 - 機械工程學系
Department of Mechanical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 139
中文關鍵詞: 視覺觸覺融合機器手臂夾取實例分割手眼校正觸覺物件辨識物件自適應夾取
外文關鍵詞: Vision-Tactile Fusion, Robotic Grasping, Instance Segmentation, Hand-Eye Calibration, Tactile Object Recognition, Object-Adaptive Grasping
相關次數: 點閱:3下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究設計並實作一套整合視覺與觸覺感測回饋之機器手臂夾取系統,鎖定雞蛋、蘋果、番茄三類食品物件,探討視觸覺融合設計於易碎與可變形物件夾取任務中之可行性。系統硬體採用達明 TM7s 六軸協作型機械手臂,搭配眼在手上(Eye-in-Hand)架構之 RealSense D435i 深度相機,並於 Robotiq 2F-85 電動夾爪指尖安裝自製 8 通道壓阻式觸覺感測陣列。視覺定位採用 YOLOv11s-seg 實例分割模型完成物件偵測與遮罩生成,以影像矩質心定位與短邊對齊決定夾取點與旋轉角度;手眼校正採用 Daniilidis 演算法,綜合空間誤差為 5.98 mm。觸覺辨識方面,系統於夾取接觸過程中分階段擷取 8 通道觸覺訊號,形成 16 維特徵向量,並以多層感知器(MLP)分類器即時判斷物件類別;訓練資料集共 90 筆(三類物件各 30 筆),獨立測試集準確率達 83.33%,5-Fold 交叉驗證平均準確率為 90.00% ± 4.16%。系統依辨識結果選取對應之夾爪出力設定值與深度補償量,實現物件自適應夾取控制。
    系統依九階段全自動夾取序列完成「全域粗估、局部精估、插入、觸覺探測與自適應夾取、拔起驗證、安全放回」之夾取流程,針對三類物件各執行 10 次實機測試,整體平均成功率達 80%(24/30),番茄最高(90%)、蘋果次之(80%)、雞蛋最低(70%)。雞蛋因觸覺辨識易與蘋果混淆而風險較高,實測結果與預期相符;番茄之實測表現則優於預期,顯示動態零點校正機制對彈性形變具一定補償效果。本研究初步驗證視覺定位、觸覺物件辨識與自適應夾取控制三者整合,有助於提升機器手臂對不同力學特性食品物件之夾取穩定性與安全性。

    Grasping food objects poses challenges beyond those of rigid industrial parts. Eggs are brittle thin-shelled bodies, tomatoes are soft and highly deformable, and apples are comparatively rigid, so no single fixed gripping strategy can simultaneously guarantee product integrity and grasp stability. Vision alone supplies geometry and pose but no contact-force feedback during closure. This study designs and implements a robotic grasping system that integrates visual guidance with tactile feedback, validated on eggs, apples, and tomatoes. The platform comprises a Techman TM7s six-axis collaborative manipulator, an Intel RealSense D435i depth camera in an eye-in-hand configuration, and a Robotiq 2F-85 electric gripper fitted with a custom eight-channel piezoresistive tactile array. A YOLOv11s-seg instance segmentation model detects the target and generates its mask; the image-moment centroid and the minimum-area rectangle determine the grasp point and the end-effector rotation. Hand-eye calibration solved by the Daniilidis dual-quaternion method attains a composite spatial error of 5.98 mm. During closure, dynamic zeroing establishes a contact-referenced depth origin, from which shallow and deep tactile features are extracted and concatenated into a sixteen-dimensional vector classified online by a multilayer perceptron. The classifier reaches 83.33% accuracy on an independent test set and 90.00% ± 4.16% under five-fold cross-validation. The recognized class then selects a corresponding gripper force register value and stroke compensation, realizing context-aware parameter switching. Over a nine-stage fully automated sequence with ten trials per object, the system achieves an overall success rate of 80% (24/30): 90% for tomatoes, 80% for apples, and 70% for eggs.

    摘要i 誌謝viii 目錄ix 圖目錄xv 表目錄xvii 第一章 緒論1 1.1 研究背景1 1.2 文獻回顧2 1.2.1 視覺導向機器人夾取技術2 1.2.2 觸覺感測與材質辨識3 1.2.3 視覺觸覺融合夾取4 1.2.4 手眼校正方法5 1.2.5 易碎與可變形物件之食品夾取6 1.3 研究目的8 1.4 論文架構9 第二章 研究理論10 2.1 視覺實例分割理論基礎10 2.1.1 物件偵測技術發展脈絡10 2.1.2 YOLO 系列網路架構原理11 2.1.3 實例分割與語意分割之原理差異12 2.1.4 物件偵測評估指標定義13 2.2 手眼校正理論基礎15 2.2.1 剛體運動與齊次座標轉換表示法15 2.2.2 眼在手上與眼在手外配置比較15 2.2.3 手眼校正問題之數學定義16 2.2.4 對偶四元數表示法原理17 2.2.5 同步求解旋轉平移之最佳化原理18 2.3 壓阻式觸覺感測理論基礎18 2.3.1 觸覺感測技術分類與原理比較19 2.3.2 壓阻效應物理原理21 2.3.3 類比訊號擷取與轉換原理22 2.3.4 訊號雜訊來源與濾波原理24 2.4 多層感知器與正則化理論基礎26 2.4.1 多層感知器基本原理27 2.4.2 反向傳播與梯度下降最佳化原理28 2.4.3 過擬合問題與正則化技術原理29 2.4.4 多類別分類之損失函數30 2.5 物件自適應策略原理31 2.5.1 物件力學特性與夾持行為之關係31 2.5.2 接觸力控制之兩種實現途徑32 2.5.3 物件類別差異對夾持參數設計之影響原理34 第三章 系統架構與硬體設計35 3.1 系統整體架構35 3.1.1 ROS 2通訊層36 3.1.2 防斷線機制38 3.2 硬體設備38 3.2.1 達明TM7S六軸協作型機器手臂39 3.2.2 Intel RealSense D435i 深度相機39 3.2.3 Robotiq 2F-85電動夾爪40 3.2.4 Raspberry Pi 4B41 3.3 觸覺感測模組設計與製作41 3.3.1 三明治疊層結構概述41 3.3.2 指尖基座3D列印設計42 3.3.3 感測層組裝與矽膠保護層44 3.3.4 ADS1115 雙模組配置與電路佈線45 3.3.5 硬體抗雜訊設計48 3.3.6 完整訊號鏈路49 3.4 觸覺感測器靜力學校正50 3.4.1 校正設備與前置處理50 3.4.2 校正流程52 3.4.3 數學模型比較與選定53 3.5 ROS 2 觸覺感測發布節點57 3.5.1 節點架構概述57 3.5.2 雙核交錯平行讀取58 3.5.3 中位數濾波與自動歸零58 3.5.4 指數模型換算與話題發布59 3.6 YOLO 實例分割模型訓練60 3.6.1 模型選用與網路架構60 3.6.2 訓練資料集建立與資料增生61 3.6.3 模型訓練設定與結果64 3.7 手眼校正實驗71 3.7.1 Eye-in-Hand 架構原理71 3.7.2 校正流程與矩陣求解71 3.7.3 演算法比較與校正結果驗證73 3.8 夾取動作與觸覺感測75 3.8.1 動態零點校正75 3.8.2 分階段力道特徵擷取與物理單調性修正77 3.8.3 MLP 推論流程80 3.8.4 微步推進探測機制81 3.9 MLP 分類器模型訓練82 3.9.1 訓練資料收集82 3.9.2 網路架構與訓練設定82 3.9.3 模型訓練結果84 第四章 視覺觸覺融合夾取系統88 4.1 視覺引導定位88 4.1.1 RGB-D 對齊與物件偵測88 4.1.2 座標轉換原理88 4.1.3 演算法設計90 4.1.4 雙階段視覺定位與距離門檻防護92 4.2 物件類別自適應夾取策略94 4.2.1 三類物件自適應策略94 4.2.2 夾取指令下達與流程收尾96 4.3 九階段全自動夾取序列96 4.3.1 到位確認機制96 4.3.2 九階段流程說明97 4.4 視覺觸覺決策層融合之離線可行性評估101 4.4.1 融合架構設計理念101 4.4.2 融合公式設計101 4.4.3 驗證方法與結果102 4.4.4 結果討論與限制104 第五章 實驗結果與討論106 5.1 實驗環境設置106 5.2 整體夾取成功率驗證106 5.3 結果討論與限制109 5.3.1 基於系統設計數據之預期表現分析109 5.3.2 三類物體之預期風險與觀察重點111 5.3.3 系統限制與未來展望銜接112 第六章 結論與未來展望114 6.1 結論114 6.2 未來展望115 參考文獻116

    [1]G. Du, K. Wang, S. Lian, and K. Zhao, "Vision-based robotic grasping from object localization, object pose estimation to grasp estimation for parallel grippers: a review," arXiv preprint arXiv:1905.06658, 2019.
    [2]T. Li, Y. Yan, C. Yu, J. An, Y. Wang, and G. Chen, "A comprehensive review of robot intelligent grasping based on tactile perception," Robotics and Computer-Integrated Manufacturing, vol. 90, p. 102792, 2024.
    [3]D. Zheng and Y. Chen, "Enhancing Robotic Grasping Detection Using Visual–Tactile Fusion Perception," Sensors, vol. 26, no. 2, p. 724, 2026.
    [4]C. Li et al., "Maniskill-vitac 2025: Challenge on manipulation skill learning with vision and tactile sensing," arXiv preprint arXiv:2411.12503, 2024.
    [5]Y. Zhu, D. Yang, and Y. Lee, "Deformable and Fragile Object Manipulation: A Review and Prospects," Sensors, vol. 25, no. 17, p. 5430, 2025.
    [6]W. Yong, X. Shunfa, and C. Konghao, "YOLOv8-LBP: multi-scale attention enhanced YOLOv8 for ripe tomato detection and harvesting keypoint localization," Frontiers in Plant Science, vol. 16, p. 1656381, 2025.
    [7]D. Song, P. Liu, Y. Zhu, T. Li, and K. Zhang, "FGS-YOLOv8s-seg: A Lightweight and Efficient Instance Segmentation Model for Detecting Tomato Maturity Levels in Greenhouse Environments," Agronomy, vol. 15, no. 7, p. 1687, 2025.
    [8]B. Yan and Q. Wu, "Visual Understanding of Intelligent Apple Picking: Detection-Segmentation Joint Architecture Based on Improved YOLOv11," Horticulturae, vol. 12, no. 4, p. 494, 2026.
    [9]F. Zhu, W. Zhang, S. Wang, B. Jiang, X. Feng, and Q. Zhao, "Apple-harvesting robot based on the YOLOv5-RACF model," Biomimetics, vol. 9, no. 8, p. 495, 2024.
    [10]T. Zhang, J. Huang, J. Niu, Z. Liu, L. Zhang, and H. Song, "Occlusion Avoidance for Harvesting Robots: A Lightweight Active Perception Model," Sensors (Basel, Switzerland), vol. 26, no. 1, p. 291, 2026.
    [11]L. Fu, Y. Majeed, X. Zhang, M. Karkee, and Q. Zhang, "Faster R–CNN–based apple detection in dense-foliage fruiting-wall trees using RGB and depth features for robotic harvesting," Biosystems Engineering, vol. 197, pp. 245–256, 2020.
    [12]I. Sa, Z. Ge, F. Dayoub, B. Upcroft, T. Perez, and C. McCool, "Deepfruits: A fruit detection system using deep neural networks," sensors, vol. 16, no. 8, p. 1222, 2016.
    [13]J. Wang and W. Sun, "Cluster segmentation and stereo vision-based apple localization algorithm for robotic harvesting," Frontiers in Plant Science, vol. 16, p. 1598414, 2025.
    [14]V. Rajendran et al., "Towards autonomous selective harvesting: A review of robot perception, robot design, motion planning and control," Journal of Field Robotics, vol. 41, no. 7, pp. 2247–2279, 2024.
    [15]Y. Tan, X. Liu, J. Zhang, Y. Wang, and Y. Hu, "A review of research on fruit and vegetable picking robots based on deep learning," Sensors, vol. 25, no. 12, p. 3677, 2025.
    [16]Y. Huang et al., "A review of visual perception technology for intelligent fruit harvesting robots," Frontiers in Plant Science, vol. 16, p. 1646871, 2025.
    [17]Y. Song, S. Lv, F. Wang, and M. Li, "Hardness-and-type recognition of different objects based on a novel porous graphene flexible tactile sensor array," Micromachines, vol. 14, no. 1, p. 217, 2023.
    [18]H. Zhou, X. Wang, H. Kang, and C. Chen, "A tactile-enabled grasping method for robotic fruit harvesting," arXiv preprint arXiv:2110.09051, 2021.
    [19]Y. Wei, L. Cai, H. Fang, and H. Chen, "Fruit recognition and classification based on tactile information of flexible hand," Sensors and Actuators A: Physical, vol. 370, p. 115224, 2024.
    [20]Z. Liao, Y. Du, J. Duan, H. Liang, and M. Y. Wang, "Quantitative Hardness Assessment with Vision-based Tactile Sensing for Fruit Classification and Grasping," arXiv preprint arXiv:2505.05725, 2025.
    [21]N. Jamali and C. Sammut, "Majority voting: Material classification by tactile sensing using surface texture," IEEE Transactions on Robotics, vol. 27, no. 3, pp. 508–521, 2011.
    [22]A. Drimus, G. Kootstra, A. Bilberg, and D. Kragic, "Design of a flexible tactile sensor for classification of rigid and deformable objects," Robotics and Autonomous Systems, vol. 62, no. 1, pp. 3–15, 2014.
    [23]F. Pastor et al., "Bayesian and neural inference on lstm-based object recognition from tactile and kinesthetic information," IEEE Robotics and Automation Letters, vol. 6, no. 1, pp. 231–238, 2020.
    [24]G. Rouhafzay and A.-M. Cretu, "An application of deep learning to tactile data for object recognition under visual guidance," Sensors, vol. 19, no. 7, p. 1534, 2019.
    [25]F. Ma, Y. Li, M. Chen, and W. Yu, "A data-driven robotic tactile material recognition system based on electrode array bionic finger sensors," Sensors and Actuators A: Physical, vol. 363, p. 114727, 2023.
    [26]S. Li et al., "Visual–tactile fusion for transparent object grasping in complex backgrounds," IEEE Transactions on Robotics, vol. 39, no. 5, pp. 3838–3856, 2023.
    [27]S. Cui, R. Wang, J. Wei, F. Li, and S. Wang, "Grasp state assessment of deformable objects using visual-tactile fusion perception," in 2020 IEEE International Conference on Robotics and Automation (ICRA), 2020: IEEE, pp. 538–544.
    [28]Z. Ding, G. Chen, Z. Wang, and L. Sun, "Adaptive visual–tactile fusion recognition for robotic operation of multi-material system," Frontiers in Neurorobotics, vol. 17, p. 1181383, 2023.
    [29]R. Wen, K. Yuan, Q. Wang, S. Heng, and Z. Li, "Force-guided high-precision grasping control of fragile and deformable objects using semg-based force prediction," IEEE Robotics and Automation Letters, vol. 5, no. 2, pp. 2762–2769, 2020.
    [30]R. Y. Tsai and R. K. Lenz, "A new technique for fully autonomous and efficient 3 d robotics hand/eye calibration," IEEE Transactions on robotics and automation, vol. 5, no. 3, pp. 345–358, 1989.
    [31]K. Daniilidis, "Hand-eye calibration using dual quaternions," The International Journal of Robotics Research, vol. 18, no. 3, pp. 286–298, 1999.
    [32]G. Li, S. Zou, S. Din, and B. Qi, "Modified hand–eye calibration using dual quaternions," Applied Sciences, vol. 12, no. 23, p. 12480, 2022.
    [33]Q. Wang et al., "Towards Damage‐Less Robotic Fragile Fruit Grasping: A Systematic Review on System Design, End Effector, and Visual and Tactile Feedback," Journal of Field Robotics, vol. 42, no. 8, pp. 4521–4543, 2025.
    [34]Y. Zhang and Z. Wang, "Review of robotic grippers for high-speed handling of fragile foods," Advanced Robotics, vol. 39, no. 17, pp. 1054–1070, 2025.
    [35]Y. Xie, B. Zhang, J. Zhou, Y. Bai, and M. Zhang, "An integrated multi-sensor network for adaptive grasping of fragile fruits: Design and feasibility tests," Sensors, vol. 20, no. 17, p. 4973, 2020.
    [36]S. Ren, K. He, R. Girshick, and J. Sun, "Faster r-cnn: Towards real-time object detection with region proposal networks," Advances in neural information processing systems, vol. 28, 2015.
    [37]J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, "You only look once: Unified, real-time object detection," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 779–788.
    [38]T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, "Feature pyramid networks for object detection," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2117–2125.
    [39]J. Terven, D.-M. Córdova-Esparza, and J.-A. Romero-González, "A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas," Machine learning and knowledge extraction, vol. 5, no. 4, pp. 1680–1716, 2023.
    [40]K. He, G. Gkioxari, P. Dollár, and R. Girshick, "Mask r-cnn," in Proceedings of the IEEE international conference on computer vision, 2017, pp. 2961–2969.
    [41]D. Bolya, C. Zhou, F. Xiao, and Y. J. Lee, "Yolact: Real-time instance segmentation," in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9157–9166.
    [42]Y. Peng, N. Yang, Q. Xu, Y. Dai, and Z. Wang, "Recent advances in flexible tactile sensors for intelligent systems," Sensors, vol. 21, no. 16, p. 5392, 2021.
    [43]Y. Shen et al., "Thin and soft optical tactile sensor for highly sensitive object perception," Optics Express, vol. 34, no. 12, pp. 21443–21458, 2026.
    [44]D. E. Rumelhart, G. E. Hinton, and R. J. Williams, "Learning representations by back-propagating errors," nature, vol. 323, no. 6088, pp. 533–536, 1986.
    [45]D. P. Kingma and J. Ba, "Adam: A method for stochastic optimization," arXiv preprint arXiv:1412.6980, 2014.
    [46]N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, "Dropout: a simple way to prevent neural networks from overfitting," The journal of machine learning research, vol. 15, no. 1, pp. 1929–1958, 2014.
    [47]N. Hogan, "Impedance control: An approach to manipulation," in 1984 American control conference, 1984: IEEE, pp. 304–313.

    QR CODE