| 研究生: |
林聿翔 Lin, Yu-Siang |
|---|---|
| 論文名稱: |
結合視觸覺之食物物件自適應智慧夾取系統研究 Research on an Adaptive and Intelligent Grasping System for Food Objects Based on Vision-Tactile Fusion |
| 指導教授: |
鍾俊輝
Chung, Chun-hui |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 機械工程學系 Department of Mechanical Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 139 |
| 中文關鍵詞: | 視覺觸覺融合 、機器手臂夾取 、實例分割 、手眼校正 、觸覺物件辨識 、物件自適應夾取 |
| 外文關鍵詞: | Vision-Tactile Fusion, Robotic Grasping, Instance Segmentation, Hand-Eye Calibration, Tactile Object Recognition, Object-Adaptive Grasping |
| 相關次數: | 點閱:3 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究設計並實作一套整合視覺與觸覺感測回饋之機器手臂夾取系統,鎖定雞蛋、蘋果、番茄三類食品物件,探討視觸覺融合設計於易碎與可變形物件夾取任務中之可行性。系統硬體採用達明 TM7s 六軸協作型機械手臂,搭配眼在手上(Eye-in-Hand)架構之 RealSense D435i 深度相機,並於 Robotiq 2F-85 電動夾爪指尖安裝自製 8 通道壓阻式觸覺感測陣列。視覺定位採用 YOLOv11s-seg 實例分割模型完成物件偵測與遮罩生成,以影像矩質心定位與短邊對齊決定夾取點與旋轉角度;手眼校正採用 Daniilidis 演算法,綜合空間誤差為 5.98 mm。觸覺辨識方面,系統於夾取接觸過程中分階段擷取 8 通道觸覺訊號,形成 16 維特徵向量,並以多層感知器(MLP)分類器即時判斷物件類別;訓練資料集共 90 筆(三類物件各 30 筆),獨立測試集準確率達 83.33%,5-Fold 交叉驗證平均準確率為 90.00% ± 4.16%。系統依辨識結果選取對應之夾爪出力設定值與深度補償量,實現物件自適應夾取控制。
系統依九階段全自動夾取序列完成「全域粗估、局部精估、插入、觸覺探測與自適應夾取、拔起驗證、安全放回」之夾取流程,針對三類物件各執行 10 次實機測試,整體平均成功率達 80%(24/30),番茄最高(90%)、蘋果次之(80%)、雞蛋最低(70%)。雞蛋因觸覺辨識易與蘋果混淆而風險較高,實測結果與預期相符;番茄之實測表現則優於預期,顯示動態零點校正機制對彈性形變具一定補償效果。本研究初步驗證視覺定位、觸覺物件辨識與自適應夾取控制三者整合,有助於提升機器手臂對不同力學特性食品物件之夾取穩定性與安全性。
Grasping food objects poses challenges beyond those of rigid industrial parts. Eggs are brittle thin-shelled bodies, tomatoes are soft and highly deformable, and apples are comparatively rigid, so no single fixed gripping strategy can simultaneously guarantee product integrity and grasp stability. Vision alone supplies geometry and pose but no contact-force feedback during closure. This study designs and implements a robotic grasping system that integrates visual guidance with tactile feedback, validated on eggs, apples, and tomatoes. The platform comprises a Techman TM7s six-axis collaborative manipulator, an Intel RealSense D435i depth camera in an eye-in-hand configuration, and a Robotiq 2F-85 electric gripper fitted with a custom eight-channel piezoresistive tactile array. A YOLOv11s-seg instance segmentation model detects the target and generates its mask; the image-moment centroid and the minimum-area rectangle determine the grasp point and the end-effector rotation. Hand-eye calibration solved by the Daniilidis dual-quaternion method attains a composite spatial error of 5.98 mm. During closure, dynamic zeroing establishes a contact-referenced depth origin, from which shallow and deep tactile features are extracted and concatenated into a sixteen-dimensional vector classified online by a multilayer perceptron. The classifier reaches 83.33% accuracy on an independent test set and 90.00% ± 4.16% under five-fold cross-validation. The recognized class then selects a corresponding gripper force register value and stroke compensation, realizing context-aware parameter switching. Over a nine-stage fully automated sequence with ten trials per object, the system achieves an overall success rate of 80% (24/30): 90% for tomatoes, 80% for apples, and 70% for eggs.
[1]G. Du, K. Wang, S. Lian, and K. Zhao, "Vision-based robotic grasping from object localization, object pose estimation to grasp estimation for parallel grippers: a review," arXiv preprint arXiv:1905.06658, 2019.
[2]T. Li, Y. Yan, C. Yu, J. An, Y. Wang, and G. Chen, "A comprehensive review of robot intelligent grasping based on tactile perception," Robotics and Computer-Integrated Manufacturing, vol. 90, p. 102792, 2024.
[3]D. Zheng and Y. Chen, "Enhancing Robotic Grasping Detection Using Visual–Tactile Fusion Perception," Sensors, vol. 26, no. 2, p. 724, 2026.
[4]C. Li et al., "Maniskill-vitac 2025: Challenge on manipulation skill learning with vision and tactile sensing," arXiv preprint arXiv:2411.12503, 2024.
[5]Y. Zhu, D. Yang, and Y. Lee, "Deformable and Fragile Object Manipulation: A Review and Prospects," Sensors, vol. 25, no. 17, p. 5430, 2025.
[6]W. Yong, X. Shunfa, and C. Konghao, "YOLOv8-LBP: multi-scale attention enhanced YOLOv8 for ripe tomato detection and harvesting keypoint localization," Frontiers in Plant Science, vol. 16, p. 1656381, 2025.
[7]D. Song, P. Liu, Y. Zhu, T. Li, and K. Zhang, "FGS-YOLOv8s-seg: A Lightweight and Efficient Instance Segmentation Model for Detecting Tomato Maturity Levels in Greenhouse Environments," Agronomy, vol. 15, no. 7, p. 1687, 2025.
[8]B. Yan and Q. Wu, "Visual Understanding of Intelligent Apple Picking: Detection-Segmentation Joint Architecture Based on Improved YOLOv11," Horticulturae, vol. 12, no. 4, p. 494, 2026.
[9]F. Zhu, W. Zhang, S. Wang, B. Jiang, X. Feng, and Q. Zhao, "Apple-harvesting robot based on the YOLOv5-RACF model," Biomimetics, vol. 9, no. 8, p. 495, 2024.
[10]T. Zhang, J. Huang, J. Niu, Z. Liu, L. Zhang, and H. Song, "Occlusion Avoidance for Harvesting Robots: A Lightweight Active Perception Model," Sensors (Basel, Switzerland), vol. 26, no. 1, p. 291, 2026.
[11]L. Fu, Y. Majeed, X. Zhang, M. Karkee, and Q. Zhang, "Faster R–CNN–based apple detection in dense-foliage fruiting-wall trees using RGB and depth features for robotic harvesting," Biosystems Engineering, vol. 197, pp. 245–256, 2020.
[12]I. Sa, Z. Ge, F. Dayoub, B. Upcroft, T. Perez, and C. McCool, "Deepfruits: A fruit detection system using deep neural networks," sensors, vol. 16, no. 8, p. 1222, 2016.
[13]J. Wang and W. Sun, "Cluster segmentation and stereo vision-based apple localization algorithm for robotic harvesting," Frontiers in Plant Science, vol. 16, p. 1598414, 2025.
[14]V. Rajendran et al., "Towards autonomous selective harvesting: A review of robot perception, robot design, motion planning and control," Journal of Field Robotics, vol. 41, no. 7, pp. 2247–2279, 2024.
[15]Y. Tan, X. Liu, J. Zhang, Y. Wang, and Y. Hu, "A review of research on fruit and vegetable picking robots based on deep learning," Sensors, vol. 25, no. 12, p. 3677, 2025.
[16]Y. Huang et al., "A review of visual perception technology for intelligent fruit harvesting robots," Frontiers in Plant Science, vol. 16, p. 1646871, 2025.
[17]Y. Song, S. Lv, F. Wang, and M. Li, "Hardness-and-type recognition of different objects based on a novel porous graphene flexible tactile sensor array," Micromachines, vol. 14, no. 1, p. 217, 2023.
[18]H. Zhou, X. Wang, H. Kang, and C. Chen, "A tactile-enabled grasping method for robotic fruit harvesting," arXiv preprint arXiv:2110.09051, 2021.
[19]Y. Wei, L. Cai, H. Fang, and H. Chen, "Fruit recognition and classification based on tactile information of flexible hand," Sensors and Actuators A: Physical, vol. 370, p. 115224, 2024.
[20]Z. Liao, Y. Du, J. Duan, H. Liang, and M. Y. Wang, "Quantitative Hardness Assessment with Vision-based Tactile Sensing for Fruit Classification and Grasping," arXiv preprint arXiv:2505.05725, 2025.
[21]N. Jamali and C. Sammut, "Majority voting: Material classification by tactile sensing using surface texture," IEEE Transactions on Robotics, vol. 27, no. 3, pp. 508–521, 2011.
[22]A. Drimus, G. Kootstra, A. Bilberg, and D. Kragic, "Design of a flexible tactile sensor for classification of rigid and deformable objects," Robotics and Autonomous Systems, vol. 62, no. 1, pp. 3–15, 2014.
[23]F. Pastor et al., "Bayesian and neural inference on lstm-based object recognition from tactile and kinesthetic information," IEEE Robotics and Automation Letters, vol. 6, no. 1, pp. 231–238, 2020.
[24]G. Rouhafzay and A.-M. Cretu, "An application of deep learning to tactile data for object recognition under visual guidance," Sensors, vol. 19, no. 7, p. 1534, 2019.
[25]F. Ma, Y. Li, M. Chen, and W. Yu, "A data-driven robotic tactile material recognition system based on electrode array bionic finger sensors," Sensors and Actuators A: Physical, vol. 363, p. 114727, 2023.
[26]S. Li et al., "Visual–tactile fusion for transparent object grasping in complex backgrounds," IEEE Transactions on Robotics, vol. 39, no. 5, pp. 3838–3856, 2023.
[27]S. Cui, R. Wang, J. Wei, F. Li, and S. Wang, "Grasp state assessment of deformable objects using visual-tactile fusion perception," in 2020 IEEE International Conference on Robotics and Automation (ICRA), 2020: IEEE, pp. 538–544.
[28]Z. Ding, G. Chen, Z. Wang, and L. Sun, "Adaptive visual–tactile fusion recognition for robotic operation of multi-material system," Frontiers in Neurorobotics, vol. 17, p. 1181383, 2023.
[29]R. Wen, K. Yuan, Q. Wang, S. Heng, and Z. Li, "Force-guided high-precision grasping control of fragile and deformable objects using semg-based force prediction," IEEE Robotics and Automation Letters, vol. 5, no. 2, pp. 2762–2769, 2020.
[30]R. Y. Tsai and R. K. Lenz, "A new technique for fully autonomous and efficient 3 d robotics hand/eye calibration," IEEE Transactions on robotics and automation, vol. 5, no. 3, pp. 345–358, 1989.
[31]K. Daniilidis, "Hand-eye calibration using dual quaternions," The International Journal of Robotics Research, vol. 18, no. 3, pp. 286–298, 1999.
[32]G. Li, S. Zou, S. Din, and B. Qi, "Modified hand–eye calibration using dual quaternions," Applied Sciences, vol. 12, no. 23, p. 12480, 2022.
[33]Q. Wang et al., "Towards Damage‐Less Robotic Fragile Fruit Grasping: A Systematic Review on System Design, End Effector, and Visual and Tactile Feedback," Journal of Field Robotics, vol. 42, no. 8, pp. 4521–4543, 2025.
[34]Y. Zhang and Z. Wang, "Review of robotic grippers for high-speed handling of fragile foods," Advanced Robotics, vol. 39, no. 17, pp. 1054–1070, 2025.
[35]Y. Xie, B. Zhang, J. Zhou, Y. Bai, and M. Zhang, "An integrated multi-sensor network for adaptive grasping of fragile fruits: Design and feasibility tests," Sensors, vol. 20, no. 17, p. 4973, 2020.
[36]S. Ren, K. He, R. Girshick, and J. Sun, "Faster r-cnn: Towards real-time object detection with region proposal networks," Advances in neural information processing systems, vol. 28, 2015.
[37]J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, "You only look once: Unified, real-time object detection," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 779–788.
[38]T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, "Feature pyramid networks for object detection," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2117–2125.
[39]J. Terven, D.-M. Córdova-Esparza, and J.-A. Romero-González, "A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas," Machine learning and knowledge extraction, vol. 5, no. 4, pp. 1680–1716, 2023.
[40]K. He, G. Gkioxari, P. Dollár, and R. Girshick, "Mask r-cnn," in Proceedings of the IEEE international conference on computer vision, 2017, pp. 2961–2969.
[41]D. Bolya, C. Zhou, F. Xiao, and Y. J. Lee, "Yolact: Real-time instance segmentation," in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9157–9166.
[42]Y. Peng, N. Yang, Q. Xu, Y. Dai, and Z. Wang, "Recent advances in flexible tactile sensors for intelligent systems," Sensors, vol. 21, no. 16, p. 5392, 2021.
[43]Y. Shen et al., "Thin and soft optical tactile sensor for highly sensitive object perception," Optics Express, vol. 34, no. 12, pp. 21443–21458, 2026.
[44]D. E. Rumelhart, G. E. Hinton, and R. J. Williams, "Learning representations by back-propagating errors," nature, vol. 323, no. 6088, pp. 533–536, 1986.
[45]D. P. Kingma and J. Ba, "Adam: A method for stochastic optimization," arXiv preprint arXiv:1412.6980, 2014.
[46]N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, "Dropout: a simple way to prevent neural networks from overfitting," The journal of machine learning research, vol. 15, no. 1, pp. 1929–1958, 2014.
[47]N. Hogan, "Impedance control: An approach to manipulation," in 1984 American control conference, 1984: IEEE, pp. 304–313.