簡易檢索 / 詳目顯示

研究生: 陳亦宥
Chen, Yi-Yu
論文名稱: 使用視覺-語言-動作流模型並透過 ArUco 微調之機器人自主導航系統
A Vision-Language-Action Flow Model for Autonomous Robot Navigation System with ArUco Refinement
指導教授: 連震杰
Lien, Jenn-Jier James
學位類別: 碩士
Master
系所名稱: 工學院 - 智慧製造國際碩士學位學程
International Master Program on Intelligent Manufacturing
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 72
中文關鍵詞: 視覺語言導航視覺-語言-動作模型流匹配導航模仿學習ArUco 標記
外文關鍵詞: Vision-Language Navigation, Vision-Language-Action, Flow Matching, Navigation, Imitation Learning, ArUco Marker
相關次數: 點閱:22下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 傳統機器人導航通常需要預先在環境中建立地圖,並依賴基於座標的路徑規劃演算法。這種僵化的依賴性限制了機器人理解抽象自然語言指令、以及適應動態或非結構化環境的能力。為了解決此限制,本論文探討了三種漸進式的先進導航方法:(1) 結合 ArUco 標記微調、達到公分級精度(平均誤差 2.15 公分)的傳統 RTAB-Map 導航系統;(2) 使用 3D 光達與 AMCL 定位的地圖式強化學習導航系統;以及 (3) 基於該架構的免建圖視覺-語言-動作導航系統。該 VLA 模型運用流匹配技術,將預訓練的視覺語言動作模型從文字生成轉用於連續的機器人動作生成,並能夠輸出高頻(50 Hz)且具時間連續性的一大段動作。透過人類遙控示範進行模仿學習訓練,此系統已在真實環境中的實體機器人上完成評估。在受控環境下進行的 20 次實驗結果顯示,系統達到 100% 的成功率與 9.50 公分的平均定位誤差,並能順利遵循如「走到放有紅色杯子的桌子旁」等抽象語言指令。此 VLA 方法無需預先建置地圖或進行明確的座標計算,即可實現流暢的導航。此外,ArUco 微調模組成功整合至傳統的 Nav2 架構(將誤差從 7.30 公分降至 2.15 公分)以及 VLA 系統中,驗證了基於視覺標記的終點精修在不同導航方法下的有效性。本研究突顯了以流模型為基礎的 VLA 模型結合 ArUco 微調,在推動自主移動機器人朝向通用、具適應性的具身智慧發展上的巨大潛力。

    Traditional robotic navigation typically requires pre-mapped environments and coordinate-based path planning algorithms. This rigid dependency limits a robot's ability to interpret abstract, natural language instructions and adapt to dynamic or unstructured settings. To address this limitation, this thesis investigates three progressively advanced navigation paradigms: (1) a classical RTAB-Map-based navigation system with ArUco refinement achieving centimeter-level precision (2.15 cm mean error), (2) a map-based reinforcement learning (RL) navigation system using 3D LiDAR and AMCL localization, and (3) a mapless Vision-Language-Action (VLA) navigation system based on the {pi }_{0} architecture. The VLA model repurposes a pre-trained Vision-Language Model (VLM) from text generation to continuous robot action generation using flow matching, capable of generating high-frequency (50 Hz), temporally continuous action chunks. Trained through imitation learning on human teleoperation demonstrations, the system is evaluated on a physical robot in real-world environments. Experimental results across 20 trials in a controlled environment demonstrate 100% success rate with a mean positioning error of 9.50 cm, successfully following abstract language instructions such as "Go to the table with red cups." The VLA approach achieves smooth navigation without requiring pre-built maps or explicit coordinate calculations. The ArUco refinement module was successfully integrated with both the classical Nav2 approach (reducing error from 7.30 cm to 2.15 cm) and the VLA system, demonstrating the effectiveness of visual marker-based terminal refinement across different navigation paradigms. This research highlights the significant potential of flow-based VLA models combined with ArUco refinement in advancing toward general, adaptable Embodied AI for autonomous mobile robots.

    Abstract II 致謝 III List of Tables VI List of Figures VII Chapter 1 Introduction 1 1.1 Motivation and Objective 1 1.2 Global Framework 2 1.3 Related Works 8 1.4 Contribution 11 Chapter 2 System Setup and Specification 13 2.1 System Setup 13 2.1.1 System Setup: Tracer Mobile Base 13 2.2 Hardware Specifications 16 Chapter 3 RTAB-Map-Based Navigation with ArUco using RGB Camera and 3D LiDAR 17 3.1 Simultaneous Localization and Mapping (SLAM) Using RTAB-Map 17 3.2 Navigation Control using ROS2::Navigation2 19 3.3 Localization - Navigation Refinement Using ArUco 20 Chapter 4 Map-Based RL Navigation using 3D LiDAR 23 4.1 Map-Based RL Navigation using 3D LiDAR 23 4.1.0 Map-Based RL Navigation using 3D LiDAR Framework 23 4.1.1 Map-Based RL Navigation using 3D LiDAR Training Framework 24 4.1.2 Map-Based RL Navigation using 3D LiDAR Inference Framework 26 4.1.3 Loss Function 29 Chapter 5 Mapless-Based π0 VLA Navigation using RGB Camera 32 5.1 Mapless-Based π0 VLA Navigation using RGB Camera 32 5.1.0 Mapless-Based π0 VLA Navigation using RGB Camera Framework 32 5.1.1 Mapless-Based π0 VLA Navigation using RGB Camera Training Framework 33 5.1.2 Mapless-Based π0 VLA Navigation using RGB Camera Inference Framework 35 5.1.3 Loss Function 40 Chapter 6 Experimental Result 44 6.1.1 Data Collection 44 6.1.2 Metrics 50 6.2.1 Experimental Result 51 Chapter 7 Conclusion and Future Work 57 7.1 Conclusion 57 7.2 Future Work 58 Reference 60

    [1] L. Beyer, A. Steiner, A.S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, and T. Unterthiner, "Paligemma: A Versatile 3B VLM for Transfer," in arXiv preprint arXiv: 2407.07726, 2024.
    [2] G. Bradski, "The OpenCV Library," Dr. Dobb's Journal: Software Tools for the Professional Programmer, Vol. 25, No. 11, pp. 120-123, 2000.
    [3] D.M. Diez, C.D. Barr, and M. Cetinkaya-Rundel, "OpenIntro Statistics," Vol. 4, Boston, MA, USA: OpenIntro, 2012.
    [4] K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L.X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky, "𝜋0: A Vision-Language-Action Flow Model for General Robot Control," in arXiv preprint arXiv: 2410.24164, 2024.
    [5] K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M.R. Equi, C. Finn, N. Fusai, M.Y. Galliker, and D. Ghosh, "$pi_0.5$: A Vision-Language-Action Model with Open-World Generalization," in Conference on Robot Learning (CoRL), 2025.
    [6] A.C. Cheng, Y.D. Ji, Z.J. Yang, X.Y. Zou, J. Kautz, E. Biyik, H.X. Yin, S.F. Liu, and X.L. Wang, "NaVILA: Legged Robot Vision-Language-Action Model for Navigation," in Robotics: Science and Systems (RSS), 2025.
    [7] C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song, "Diffusion Policy: Visuomotor Policy Learning via Action Diffusion," The International Journal of Robotics Research, Vol. 44, No. 10-11, pp. 1684-1704, 2025.
    [8] C. Glossop, W. Chen, A. Bhorkar, D. Shah, and S. Levine, "CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models," in arXiv preprint, 2025.
    [9] S. Garrido-Jurado, R. Muñoz-Salinas, F.J. Madrid-Cuevas, and M.J. Marín-Jiménez, "Automatic Generation and Detection of Highly Reliable Fiducial Markers under 60 Occlusion," Pattern Recognition, Vol. 47, No. 6, pp. 2280-2292, 2014.
    [10] O.Y. Goba, A.Y. Gado, C.M. Elias, and A. Hussein, "From Prompts to Pavement: LMMs-Based Agentic Behavior-Tree Generation Framework for Autonomous Vehicles," in 2025 IEEE 28th International Conference on Intelligent Transportation Systems (ITSC), pp. 1637-1643, 2025.
    [11] H.V. Hasselt, A. Guez, and D. Silver, "Deep Reinforcement Learning with Double QLearning," in Proceedings of the AAAI Conference on Artificial Intelligence, pp. 20942100, 2016.
    [12] Y. Lipman, R.T. Chen, H. Ben-Hamu, M. Nickel, and M. Le, "Flow Matching for Generative Modeling," in arXiv preprint arXiv: 2210.02747, 2022.
    [13] Y. Lipman, M. Havasi, P. Holderrieth, N. Shaul, M. Le, B. Karrer, R.T.Q. Chen, D. Lopez-Paz, H. Ben-Hamu, and I. Gat, "Flow Matching Guide and Code," in arXiv preprint arXiv: 2412.06264, 2024.
    [14] M. Labbé and F. Michaud, "RTAB-Map as an Open-Source Lidar and Visual Simultaneous Localization and Mapping Library for Large-Scale and Long-Term Online Operation," Journal of Field Robotics, Vol. 36, No. 2, pp. 416-446, 2019.
    [15] C.J. Lin, C.C. Peng, and S.Y. Lu, "Real-Time Localization for an AMR Based on RTAB-Map," Actuators, Vol. 14, No. 3, p. 117, 2025.
    [16] S. Macenski, F. Martín, R. White, and J. Clavero, "The Marathon 2: A Navigation System," in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2020.
    [17] E. Marchesini and A. Farinelli, "Discrete Deep Reinforcement Learning for Mapless Navigation," in IEEE International Conference on Robotics and Automation (ICRA), pp. 10688-10694, 2020.
    [18] V. Mnih, A.P. Badia, M. Mirza, A. Graves, T. Harley, T.P. Lillicrap, D. Silver, and K. Kavukcuoglu, "Asynchronous Methods for Deep Reinforcement Learning," in International Conference on Machine Learning (ICML), pp. 1928-1937, 2016.
    [19] M.J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. 61 Foster, G. Lam, P. Sanketi, and Q. Vuong, "OpenVLA: An Open-Source Vision-Language-Action Model," in arXiv preprint arXiv: 2406.09246, 2024.
    [20] M.J. Kim, C. Finn, and P. Liang, "Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success," in arXiv preprint arXiv: 2502.19645, 2025.
    [21] Y. Qu, Z. Huang, Z. Sheng, J. Chen, Y. Leng, S. Labi, and S. Chen, "VLM-SAFE: Vision-Language Model-Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving," in arXiv preprint arXiv: 2505.16377, 2025.
    [22] J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, "RoFormer: Enhanced Transformer with Rotary Position Embedding," Neurocomputing, Vol. 568, p. 127063, 2024.
    [23] H. Surmann, C. Jestel, R. Marchel, F. Musberg, H. Elhadj, and M. Ardani, "Deep Reinforcement Learning for Real Autonomous Mobile Robot Navigation in Indoor Environments," in arXiv preprint arXiv: 2005.13857, 2020.
    [24] R.S. Sutton and A.G. Barto, "Reinforcement Learning: An Introduction," Cambridge, MA: MIT Press, 1998.
    [25] L. Tai, G. Paolo, and M. Liu, "Virtual-to-real Deep Reinforcement Learning: Continuous Control of Mobile Robots for Mapless Navigation," in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 31-36, 2017.
    [26] Z. Wang and X. Zhao, "Flow Matching vs. Denoising Diffusion Models: A Unified Perspective on Generative Modeling," in 2025 IEEE 4th Industrial Electronics Society Annual On-Line Conference (ONCON), pp. 1-6, 2025.
    [27] X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, "Sigmoid Loss for Language Image Pre-Training," in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 11975-11986, 2023.

    下載圖示
    校外:立即公開
    QR CODE