| 研究生: |
黑崇瑜 Hei, Chung-Yu |
|---|---|
| 論文名稱: |
應用於動態場景穩態建模的場景驅動之視覺里程計 Scenario-driven Visual Odometry and its Application to Steady-state 3D Model Reconstruction in Dynamic Scenes |
| 指導教授: |
謝明得
Shieh, Ming-Der |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2023 |
| 畢業學年度: | 111 |
| 語文別: | 英文 |
| 論文頁數: | 48 |
| 中文關鍵詞: | 同步定位與地圖構建 、三維重建 、視覺里程計 、迭代最近點演算法 、擴增實境 |
| 外文關鍵詞: | Simultaneous Localization and Mapping (SLAM), 3D reconstruction, Visual Odometry, Iterative Closest Point (ICP), Augmented Reality |
| 相關次數: | 點閱:124 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
RGBD SLAM(同時定位與建圖)是一種利用RGB-D傳感器實現的技術,可以在實時建圖的同時估計移動代理的姿態。RGB-D傳感器,如Microsoft Kinect或Intel RealSense相機,提供彩色(RGB)和深度(D)信息,相較於傳統的僅具有RGB信息的相機,能夠實現更強大和準確的感知。
RGBD SLAM在機器人技術、擴增實境、自主導航和三維重建等領域中具有廣泛應用。在機器人技術方面,它使得機器人能夠在構建空間認識的同時導航並與環境進行交互。擴增實境系統利用RGBD SLAM將虛擬內容覆蓋到現實世界中,提供沉浸式的用戶體驗。自主導航的車輛可以借助RGBD SLAM在動態環境中進行障礙物檢測、建圖和定位。RGBD SLAM在三維重建方面也可以創建室內和室外場景的詳細三維模型。
因此,我們提出了一個基於迭代最近點(Iterative Closest Point,ICP)算法的視覺里程計(VO)系統。我們的系統可以根據觀察到的RGB-D信息調整估計參數。此外,我們將該算法集成到著名的SLAM系統ElasticFusion[1]中,以創建密集的室內3D重建模型。
考慮到動態場景和實際應用,我們將輕量級的卷積神經網絡模型YOLOv7-tiny[2]與RGB幀結合使用,以檢測移動物體。我們結合深度信息創建動態物體過濾器,使我們的VO系統不受移動物體的干擾,實現僅包含靜態物體的模型重建。為了進一步提升系統性能,我們制訂了建模規則,將模型更新為當前外觀,從場景中消除移動物體,獲得最終穩態的重建結果。
我們提出的基於ICP算法、集成到ElasticFusion並通過YOLOv7-tiny模型增強的VO系統應對動態場景的挑戰。該系統過濾移動物體並相應地更新模型,實現準確的室內靜態物體3D重建。未來的研究將進一步優化和完善該系統,以實現更廣泛的實際應用。
RGBD SLAM (Simultaneous Localization and Mapping) is a technique that leverages RGB-D sensors to enable real-time mapping of the environment while simultaneously estimating the pose of a mobile agent within it. RGB-D sensors, such as Microsoft Kinect or Intel RealSense cameras, provide both color (RGB) and depth (D) information, allowing for more robust and accurate perception in comparison to traditional RGB-only cameras.
RGBD SLAM has numerous applications in robotics, augmented reality, autonomous navigation, and 3D reconstruction. In robotics, it enables robots to navigate and interact with the environment while building a spatial understanding of their surroundings. Augmented reality systems utilize RGBD SLAM to overlay virtual content onto the real world, providing immersive user experiences. Autonomous vehicles can benefit from it for obstacle detection, mapping, and localization in dynamic environments. Additionally, RGBD SLAM has applications in 3D reconstruction for creating detailed 3D models of indoor and outdoor scenes.
Therefore, we propose a Visual Odometry (VO) approach based on the Iterative Closest Point (ICP) algorithm. Our system can adjust the estimation parameters based on the observed RGB-D information. Additionally, we integrate this algorithm into the well-known SLAM system, ElasticFusion[1], to create a dense indoor 3D reconstruction.
Considering dynamic scenes and practical applications, we combine a lightweight CNN model, YOLOv7-tiny [2], to detect moving objects from RGB frames and combining depth information to create a dynamic object filter, freeing our VO system from disturbances of moving objects and enabling the reconstruction of a model that only includes static objects. We also design a model fusion rule to update our model with the current appearance, eliminating the moving objects from the scene and obtaining a final steady-state reconstruction result.
Our proposed VO system based on the ICP algorithm, integrated into ElasticFusion, and enhanced with the YOLOv7-tiny model, addresses the challenges posed by dynamic scenes. This system enables accurate indoor 3D reconstruction by filtering out moving objects and updating the model accordingly. Future research aims to further optimize and refine the system for broader practical applications.
[1] T. Whelan, R. F. Salas-Moreno, B. Glocker, A. J. Davison and S. Leutenegger, EasticFusion: Real-Time Dense SLAM and Light Source Estimation. IJRR '16
[2] C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, “YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,” 2022, arXiv:2207.02696.
[3] Hartley, R.; Zisserman, A. Multiple View Geometry in Computer Vision, 2nd ed.; Cambridge University Press: Cambridge, UK, 2004.
[4] J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cremers, “A benchmark for the evaluation of RGB-D SLAM systems,” in Proc. IEEE/RSJ Int.Conf. Intell. Robots Syst., Oct. 2012, pp. 573–580
[5] Meriam, J. L., & Kraige, L. G. (2012). Engineering mechanics: dynamics (7th ed.). John Wiley & Sons.
[6] Y. Ma, S. Soatto, J. Kosecka, and S. S. Sastry, An Invitation to 3-D Vision: From Images to Geometric Models. New York, NY, USA:Springer, 2005.
[7] Besl, P. J., & McKay, N. D. (1992). A Method for Registration of 3-D Shapes. IEEE Transactions on Pattern Analysis and Machine Intelligence, 14(2), 239-256.
[8] Kok-Lim Low. Linear least-squares optimization for point-to-plane icp surface registration. Chapel Hill, University of North Carolina, 4(10):1–3, 2004.
[9] Choi, C., A.J. Trevor, and H.I. Christensen. RGB-D edge detection and edge-based registration. in 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems. 2013. IEEE.
[10] Kanopoulos, N., Vasanthavada, N., & Baker, R. L. (1988). Design of an image edge detection filter using the Sobel operator. IEEE Journal of Solid-State Circuits, 23(2), 358–367.
[11] Canny, J., 1986. A computational approach to edge detection. IEEE Transactions on pattern analysis and machine intelligence, (6), pp.679–698.
[12] William H. Press, Saul A. Teukolsky, William T. Vetterling, and Brian P. Flannery. Numerical Recipes in C:The Art of Scientific Computing, Second Edition, Cambridge University Press, 1992.
[13] C. -W. Chen, J. Wang and M. -D. Shieh, "Edge-Based Meta-ICP Algorithm for Reliable Camera Pose Estimation," in IEEE Access, vol. 9, pp. 89020-89028, 2021, doi: 10.1109/ACCESS.2021.3090170.
[14] C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,”in Proc. Int. Conf. Mach. Learn, vol. 70, Aug. 2017, pp. 1126–1135.
[15] M. Keller, D. Lefloch, M. Lambers, S. Izadi, T. Weyrich, and A. Kolb. Real-time 3D Reconstruction in Dynamic Scenes using Point-based Fusion. In Proc. of Joint 3DIM/3DPVT Conference (3DV), 2013.
[16] Kanopoulos, N., Vasanthavada, N. & Baker, R.L., 1988. Design of an image edge detection filter using the Sobel operator. IEEE Journal of solid-state circuits, 23(2), pp.358–367.
[17] Nguyen, C.V., S. Izadi, and D. Lovell. Modeling kinect sensor noise for improved 3d reconstruction and tracking. in 2012 second international conference on 3D imaging, modeling, processing, visualization & transmission. 2012. IEEE.
[18] Li, J.-Y., Edge-aware Sampling Scheme and Sliding Estimation for ICP-based Visual Odometry, in Department of Electrical Engineering. 2020, NCKU. p.57.
[19] Tsung-Yi Lin, Maire, M., Belongie, S. J., Bourdev, L. D., Girshick, R. B., Hays, J., … Zitnick, C. L. (2014). Microsoft COCO: Common Objects in Context. CoRR, abs/1405.0312. Retrieved from http://arxiv.org/abs/1405.0312
[20] Horn, B.K., Closed-form solution of absolute orientation using unit quaternions. Josa a, 1987. 4(4): p. 629-642.
[21] R. Scona, M. Jaimez, Y. R. Petillot, M. Fallon, and D. Cremers, “StaticFusion: Background reconstruction for dense RGB-D SLAM in dynamic environments,” in Proc. IEEE Int. Conf. Robot. Autom. (ICRA), Brisbane, QLD, Australia, May 2018, pp. 1–9.
[22] R. Mur-Artal and J. D. Tardos, “ORB-SLAM2: An open-source SLAM system for monocular, stereo, and RGB-D cameras,” IEEE Trans. Robot., vol. 33, no. 5, pp. 1255–1262, Oct. 2017.
[23] R. A. Newcombe, A. J. Davison, S. Izadi, P. Kohli, O. Hilliges, J. Shotton, D. Molyneaux, S. Hodges, D. Kim, and A. Fitzgibbon, “KinectFusion: Real-time dense surface mapping and tracking,” in Proc. 10th IEEE Int. Symp. Mixed Augmented Reality, Oct. 2011, pp. 127–136
[24] Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM 65, 1 (2021), 99–106
[25] R. A. Newcombe, D. Fox and S. M. Seitz, "DynamicFusion: Reconstruction and tracking of non-rigid scenes in real-time," 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 2015, pp. 343-352, doi: 10.1109/CVPR.2015.7298631.
[26] Aljaz Bozic, Michael Zollhofer, Christian Theobalt, and Matthias Nießner. Deepdeform: Learning non-rigid rgb-d reconstruction with semi-supervised data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7002–7012, 2020
[27] Mur-Artal, R., J.M.M. Montiel, and J.D. Tardos, ORB-SLAM: a versatile and accurate monocular SLAM system. IEEE transactions on robotics, 2015. 31(5): p. 1147-1163.
[28] A. J. Davison, I. D. Reid, N. D. Molton and O. Stasse, "MonoSLAM: Real-Time Single Camera SLAM," in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 29, no. 6, pp. 1052-1067, June 2007, doi: 10.1109/TPAMI.2007.1049.