簡易檢索 / 詳目顯示

研究生: 陳俊維
Chen, Chun-Wei
論文名稱: 應用於同步定位與地圖構建系統之穩健視覺里程計演算法設計
Development of Robust Visual Odometry Algorithms for SLAM Systems
指導教授: 謝明得
Shieh, Ming-Der
學位類別: 博士
Doctor
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2021
畢業學年度: 109
語文別: 英文
論文頁數: 92
中文關鍵詞: 同步定位與地圖構建 、視覺里程計 、相機姿態估測 、迭代最近點 、元學習
外文關鍵詞: Simultaneous localization and mapping, Visual odometry, Camera pose estimation, Iterative closest point, Meta-learning
相關次數: 點閱:225  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 由於在許多應用中皆對高清3維地圖有強烈的需求,同步定位與地圖構建系統(simultaneous localization and mapping)近來受到更多的重視。視覺里程計(visual odometry)演算法作為系統中的定位模組,更是尤為關鍵。此外,其相機姿態估測的穩健度更是會輕易影響到整個系統的穩定度。本論文提出數種策略來提升姿態估測的穩健度,其中包含了例外處理機制,可靠區域偵測,考量相機速度之權重函數,元學習能力,以及具資源感知之姿態融合方法。
    為了能處理高速移動中的相機姿態估計問題,本文提出一個中斷機制以避免輸出因失真深度影像而導致的錯誤姿態估測結果。此外,本文並提出一快速關鍵幀(keyframe)選擇與切換機制以捨棄不適合之關鍵幀,並以合格預備幀替換之。所需付出之額外運算可透過在迭代最近點(iterative closest point)方法中以首次迭代產生之群內資料資訊的快速判斷來大幅降低。本文另提出了參考時間與空間上之群內資料關聯來評估關鍵幀與預備幀之品質。實驗結果顯示本方法可有效地同時降低相對姿態誤差以及運算時間。
    為了能萃取穩定區域並保有足夠的幾何結構,本文提出了一種高效近邊緣區域之噪聲區域模型。通過跳過邊緣區域附近的噪聲區域,可以有效地檢測位於邊緣附近的可靠區域。為了進一步考慮相機運動速度以進行高品質點選取,本文提出了一種運動感知採樣權重模型。並以該權重模型來估計高品質點最可能出現之深度值。最後,為了同時選取高品質採樣點並確保其分佈適當,本文採用了具適應性之非極值抑制(non-maximum suppression)方法來進行採樣。
    啟發自元學習演算法,本文提出一個元迭代最近點演算法以從多項任務資料點中共同預測其最佳姿態轉換結果。為了提升任務資料點採集效率,本文提出一基於邊緣點之任務集分割演算法以構建互補任務集。此外,為了避免部分任務卡在區域最佳解,本文採用一動態模型適配機制以擾動其任務。實驗結果顯現本方法可有效降低可能的嚴重錯誤估測,因此在重複抽樣的實驗中能取得更小的誤差分布。
    為能在困難條件下可靠地預測相機姿態,本文提出一個採用了姿態融合策略之新穎迭代最近點演算法。然而,姿態融合需重覆估計多次姿態,若使用有限運算資源之系統運行時所需耗費的時間將超出即時運算的時間限制。為了使其能達到即時姿態估測,本文提出一動態姿態融合機制以及其配套之資源配置模組。實驗結果顯示在相機高速移動之影片中,幀與幀之間之最大估測誤差能被有效的壓制。

    With a strong demand for the high quality 3D maps in many applications, simultaneous localization and mapping (SLAM) systems, which can consistently estimate the camera trajectory while simultaneously building up the surrounding map, have gained increasing attention recently. Visual odometry (VO) algorithms, which are responsible for estimating camera poses to construct the corresponding camera trajectory, play an important role in SLAM systems. In addition, the pose estimation reliability can easily influence the stability of a SLAM system. This work improves the pose estimation reliability through several strategies, including: fail-safe exception handling mechanisms, reliable region extraction, a motion-aware weighting model, the meta-learning ability, and a resource-aware pose aggregation methodology.
    To handle cases of large camera motion, an early termination mechanism is introduced to avoid producing spurious estimations caused by distorted depths. Furthermore, to replace unsuitable keyframes with qualified backup frames, a fast keyframe selection and switching algorithm is presented. The overhead of using backup process is greatly reduced by only inspecting the inlier information produced at the first iteration of iterative closest point (ICP) algorithm. Moreover, several useful criteria considering spatial and/or temporal relationships are also presented to evaluate the quality of keyframes and backup frames. By adopting the proposed schemes, both the relative pose error and the ICP iterations can be effectively reduced.
    To effectively identify regions which are reliable while preserve enough geometric structures, an efficient noisy region model for near-edge regions is presented. The reliable regions located near edges can be effectively detected by skipping the corresponding noisy regions. To further taking camera motion into account for salient point sampling, a motion-aware weighting model was presented. The optimal depth value, which covers salient points with high probability, can be estimated using the proposed weighting model. Finally, to sample salient points considering the distribution of the sampled points, an adaptive macro-block based NMS method is adopted.
    Inspired by meta-learning algorithms, this work proposes a meta-ICP algorithm to jointly estimate the optimal transformation for multiple tasks, which are constructed by sampled datapoints. To increase task sampling efficiency, an edge-based task set partition algorithm is introduced for constructing complementary task sets. In addition, to prevent the ICP from being trapped in local minima, a dynamic model adaptation scheme is adopted to disturb the trapped tasks. The experimental results reveal that the probability of unstable estimations can be effectively reduced, indicating a much narrower error distribution in repeated experiments when adopting re-sampled points.
    To reliably estimate camera poses in challenging conditions, a novel ICP adopting the pose aggregation strategy is presented. However, estimating multiple poses, which is required for performing pose aggregation, will exceed the real-time constraint when using systems with limited computation resources. To achieve real-time pose estimation for systems with limited resources, an adaptive pose aggregation scheme along with a point resource allocation module are presented. Experimental results reveal that the maximum error of frame-frame estimation can be dramatically reduced for high-speed sequences.

    摘要 I ABSTRACT III 誌謝 VI CONTENTS VII LIST OF FIGURES X LIST OF TABLES XIII 1. INTRODUCTION 1 1.1 SLAM Applications 1 1.2 Motivation 2 1.2.1 Motivation of Large Motion Handling for ICP-based Visual Odometry 2 1.2.2 Motivation of Motion-aware Keypoint Selection System for ICP-based Visual Odometry 2 1.2.3 Motivation of Edge-based Meta-ICP Algorithm for Reliable Camera Pose Estimation 3 1.2.4 Motivation of Resource-aware Edge-based ICP Algorithm for Reliable Camera Pose Estimation 4 1.3 Organization of the Dissertation 5 2. BACKGROUND 6 2.1 Coordinate Systems 6 2.1.1 Camera Model 6 2.1.2 Coordinate Transformation 7 2.1.3 Normal Estimation 8 2.2 ICP Algorithm 10 2.2.1 Correspondence Pairing 10 2.2.2 Outlier Rejection 11 2.2.3 Correspondence Weighting 11 2.2.4 Transformation Estimation 12 2.3 Control Point Selection 14 2.3.1 Geometric Stability 14 2.3.2 Depth Edge Detection 16 2.4 MAML Algorithm 18 2.5 SLAM Systems 19 2.5.1 Visual Odometry 19 2.5.2 Pose Graph Optimization 22 3. LARGE MOTION HANDLING FOR ICP-BASED VISUAL ODOMETRY 23 3.1 Motion-aware Frame Skipping Mechanism 23 3.2 Fast Keyframe Selection and Switching Strategy 27 3.2.1 Source Frame Selection 27 3.2.2 Keyframe and Backup Frame Update 29 3.2.3 Experimental Results 30 3.3 Summary 34 4. MOTION-AWARE KEYPOINT SELECTION SYSTEM FOR ICP-BASED VISUAL ODOMETRY 35 4.1 Region of Interest Detection 36 4.2 Motion-aware Correspondence Quality Modeling 41 4.3 Experimental Results 45 4.4 Summary 48 5. EDGE-BASED META-ICP ALGORITHM FOR RELIABLE CAMERA POSE ESTIMATION 49 5.1 Meta-ICP Overview 49 5.2 Edge-based Task Set Partition Algorithm 50 5.3 Dynamic Model Adaptation 55 5.4 Experimental Results 58 5.4.1 Edge Feature Comparisons 58 5.4.2 Meta-ICP Evaluation 60 5.5 Summary 65 6. RESOURCE-AWARE EDGE-BASED ICP ALGORITHM FOR RELIABLE CAMERA POSE ESTIMATION 66 6.1 Resource-aware ICP Overview 66 6.2 Adaptive Aggregation Methodology 69 6.2.1 Pose Aggregation Strategy 69 6.2.2 Spurious Probability Modeling 70 6.2.3 Model Initialization and Efficient Bagging Strategy 72 6.3 Budget-aware Resource Allocation Scheme 74 6.4 Experimental Results 77 6.5 Summary 83 CHAPTER 7 84 7. CONCLUSION AND FUTURE WORK 84 7.1 Conclusion 84 7.2 Future Work 86 BIBLIOGRAPHY 87 PUBLICATION LIST 91

    [1] P. J. Besl and N. D. McKay, “A method for registration of 3-D shapes,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 14, no. 2, pp. 239-256, Feb. 1992.
    [2] N. Gelfand, L. Ikemoto, S. Rusinkiewicz, and M. Levoy, “Geometrically stable sampling for the ICP algorithm,” in Proc. Int. Conf. 3D Digital Imaging and Modeling (3DIM), Oct. 2003.
    [3] Y. Chen and G. Medioni., “Object modeling by registration of 3-D shapes,” in Proc. IEEE Int. Conf. Robot. Autom., April 1991.
    [4] S. Rusinkiewicz and M. Levoy, “Efficient variants of the ICP algorithm,” in Proc. Int. Conf. 3D Digital Imaging and Modeling (3DIM), May 2001.
    [5] D. Simon, Fast and accurate shape-based registration, Ph.D Dissertation. Carnegie Mellon University, 1996.
    [6] C. Kerl, J. Sturm, and D. Cremers, “Dense visual SLAM for RGB-D cameras,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Nov. 2013.
    [7] C. Choi, A. J. B. Trevor, and H. I. Christensen, “RGB-D edge detection and edge-based registration,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Nov. 2013.
    [8] L. Bose and A. Richards, “Fast depth edge detection and edge based RGB-D SLAM,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), pp. 1323-1330, May 2016.
    [9] C. V. Nguyen, S. Izadi, and D. Lovell, "Modeling Kinect sensor noise for improved 3D reconstruction and tracking, " in Proc. Int. Conf. 3D Imaging, Modeling, Processing, Visualization and Transmission (3DIMPVT), Oct. 2012.
    [10] J. J. Tarrio and S. Pedre, “Realtime edge-based visual odometry for a monocular camera,” in Proc. Int. Conf. Computer Vision (ICCV), pp. 702-710, 2015.
    [11] J. Stückler and S. Behnke, “Integrating depth and color cues for dense multi-resolution scene mapping using rgb-d cameras,” IEEE Int. Conf. Multisensor Fusion and Information Integration (MFI), 2012.
    [12] M. P. Kuse and S. Shen, “Robust camera motion estimation using direct edge alignment and sub-gradient method,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), pp. 573-579, 2016.
    [13] F. Schenk and F. Fraundorfer, “Robust edge-based visual odometry using machine learned edges,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), pp. 1297-1304, 2017.
    [14] P. Dollár and C. L. Zitnick, “Fast edge detection using structured forests,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 37, no. 8, pp. 1558-1570, 2015.
    [15] S. Xie and Z. Tu, “Holistically-nested edge detection,” in Proc. Int. Conf. Computer Vision (ICCV), 2015.
    [16] J. Canny, “A computational approach to edge detection,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 8, no. 6, pp. 679-698, 1986.
    [17] K. Pulli, “Multiview registration for large data sets,” in Proc. Int. Conf. 3D Digital Imaging and Modeling (3DIM), 1999.
    [18] K. Khoshelham and S. Elberink, “Accuracy and resolution of kinect depth data for indoor mapping applications,” Sensors, vol. 12, no. 2, pp. 1437-1454, 2012.
    [19] C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in Proc. Int. Conf. Mach. Learn, vol. 70, Aug. 2017.
    [20] L. Breiman, “Bagging predictors,” in Proc. Mach. Learn., vol. 24, pp. 123-140, 1996.
    [21] J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cremers, “A benchmark for the evaluation of RGB-D SLAM systems,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Oct. 2012.
    [22] R. Mur-Artal and J. D. Tardós, “ORB-SLAM2: an open-source SLAM system for monocular, stereo and RGB-D cameras,” IEEE Trans. on Robotics, vol. 33, no. 5, pp. 1255-1262, 2017.
    [23] M. Meilland, A. Comport, and P. Rives, “Dense visual mapping of large scale environments for real-time localisation,” IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Dec. 2011.
    [24] K. Low, “Linear least-squares optimization for point-to-plane icp surface registration,” Technical report, TR04-004, University of North Carolina, 2004.
    [25] Y. Zhou, H. Li, and Kneip, “Canny-vo: Visual odometry with rgb-d cameras based on geometric 3-D-2-D edge alignment,” IEEE Trans. on Robotics, vol. 35, no. 1, pp. 184-199, 2019.
    [26] C. Kim, J. Kim, and H. J. Kim, “Edge-based visual odometry with stereo cameras using multiple oriented quadtrees,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Oct. 2020.
    [27] R. L. Gutierrez and M. Leonetti, “Information-theoretic task selection for meta-reinforcement learning,” in NeurIPS, 2020.
    [28] R. B. Rusu, “Semantic 3D object maps for everyday manipulation in human living environments,” Ph.D. dissertation, Computer Science department, Technische Universitaet Muenchen, Oct. 2009.
    [29] P. Henry, M. Krainin, E. Herbst, X. Ren, and D. Fox, “RGB-D mapping: using depth cameras for dense 3D modeling of indoor environments,” Int. Symp. on Experimental Robotics (ISER), 2010.
    [30] A. Handa, T. Whelan, J. McDonald, and A. Davison, “A benchmark for RGB-D visual odometry, 3D reconstruction and SLAM” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), pp. 1524-1531, 2014.
    [31] G. Grisetti, R. Kummerle, C. Stachniss, and W. Burgard, “A tutorial ¨ on graph-based SLAM,” IEEE Intell. Transp. Syst. Mag., vol. 2, no. 4, pp. 31–43, 2010.
    [32] C. Cadena et al., “Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age,” IEEE Trans. Robot., vol. 32, no. 6, pp. 1309–1332, Dec. 2016.
    [33] R. Mur-Artal, J. M. M. Montiel, and J. D. Tardós, “ORB-SLAM: a versatile and accurate monocular SLAM system,” IEEE Trans. Robot., vol. 31, no. 5, pp. 1147–1163, 2015.
    [34] E. Rosten and T. Drummond, “Machine learning for high-speed corner detection,” in Proc. Eur. Conf. Comput. Vision, pp. 430–443, May 2006.
    [35] R. B. Rusu and S. Cousins, “3D is here: Point cloud library,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), pp. 1–4, 2011.
    [36] Y. Zhong, “Intrinsic shape signatures: A shape descriptor for 3D object recognition,” in Proc. IEEE Int. Conf. Comput. Vis. Workshops, pp. 689–696, 2009.
    [37] D. Lowe, “Distinctive Image Features from Scale-Invariant Keypoints,” Int’l J. Computer Vision, vol. 60, no. 2, pp. 91-110, Nov. 2004.
    [38] I. Sipiran and B. Bustos, “Harris 3D: A robust extension of the Harris operator for interest point detection on 3D meshes,” The Vis. Comput.: Int. J. Comput. Graph, vol. 27, pp. 963–976, 2011.
    [39] B. Steder, R. B. Rusu, K. Konolige, and W. Burgard, “NARF: 3D range image features for object recognition,” in Proc. Workshop Defining Solving Realistic Percept. Problems Personal Robot. IEEE/ RSJ Int. Conf. Intell. Robots Syst., vol. 44, 2010.
    [40] S. Tourani, M. Sudhanshu, A. Nagariya, V. Chari, and M. Krishna, “Rolling shutter and motion blur removal for depth cameras,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), pp. 5098–5105, 2016.
    [41] A. Angeli, D. Filliat, S. Doncieux, and J. Meyer, “Fast and incremental method for loop-closure detection using bags of visual words,” IEEE Transactions on Robotics, vol. 24, no. 5, pp. 1027–1037, 2008.

    下載圖示
    2026-09-01公開
    QR CODE