簡易檢索 / 詳目顯示

研究生: 劉亭萱
Liu, Ting-Hsuan
論文名稱: 基於語意加權對比學習之跨域視覺定位研究
Semantic-Weighted Contrastive Learning for Cross-Domain Visual Localization
指導教授: 呂學展
Lu, Hsueh-Chan
學位類別: 碩士
Master
系所名稱: 工學院 - 測量及空間資訊學系
Department of Geomatics
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 87
中文關鍵詞: 視覺定位影像檢索領域自適應語意引導表徵學習對比學習
外文關鍵詞: Visual Localization, Image Retrieval, Domain Adaptation, Semantic-Guided Representation Learning, Contrastive Learning
相關次數: 點閱:47下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 隨著自動駕駛、行動機器人與擴增實境等應用快速發展,如何在長期動態環境中維持穩定且準確的視覺定位能力,已成為重要研究課題。然而,相同地點可能因天氣、光照、季節、植被與日夜變化而產生顯著外觀差異,使影像檢索式定位面臨跨域泛化困難。既有方法常透過語意分割、深度估計、領域自適應或三元損失提升定位表現,但部分方法需依賴深度資訊、時序約束或額外幾何驗證,且三元損失之效果亦容易受到樣本組合設計影響。為解決上述問題,本文提出 SeCoLoc,一個適用於長期動態環境的檢索式跨域視覺定位框架,整合語意分割、領域對抗學習與可學習語意加權對比模組(Learnable Semantic-Weighted Contrastive Module, LSWCM)。其中,LSWCM 利用語意分割結果產生 soft mask,並透過可學習語意群組權重引導特徵聚合,使模型能自動調整不同語意區域對定位任務的重要性;同時,本文根據資料集中相同位置但不同環境變化之影像建立正樣本關係,以對比學習學習跨域穩定表徵,降低人工設計三元組樣本的需求。實驗結果顯示,本文方法在 Extended CMU-Seasons 與 RobotCar-Seasons 資料集上皆展現具競爭力的定位表現,尤其在 RobotCar-Seasons 的日夜光照變化實驗中,於粗精度評估下取得最佳結果。整體而言,本文方法在不依賴深度資訊的條件下,仍能達到與現有先進方法相近甚至部分設定下更佳的定位表現,展現良好的跨域泛化能力與實務部署潛力。

    With the rapid development of autonomous driving, mobile robotics, and augmented reality, maintaining robust and accurate visual localization in long-term dynamic environments has become an important research topic. However, images captured at the same location may exhibit significant appearance changes due to weather, illumination, seasonal variations, vegetation conditions, and day-night transitions, making retrieval-based visual localization challenging under cross-domain scenarios. Existing methods often employ semantic segmentation, depth estimation, domain adaptation, or triplet loss to improve localization performance. Nevertheless, some approaches rely on depth information, temporal constraints, or additional geometric verification, while the effectiveness of triplet loss is also sensitive to manually designed sample construction. To address these issues, this study proposes SeCoLoc, a retrieval-based cross-domain visual localization framework for long-term dynamic environments, integrating semantic segmentation, adversarial domain learning, and the Learnable Semantic-Weighted Contrastive Module (LSWCM). As the core module of SeCoLoc, LSWCM generates soft masks from semantic segmentation results and uses learnable semantic group weights to guide feature aggregation, allowing the model to adaptively adjust the importance of different semantic regions for localization. In addition, instead of constructing positive samples through conventional data augmentation, this study defines positive pairs based on images captured at the same location under different environmental conditions and employs contrastive learning to learn domain-stable representations, reducing the need for manually designed triplet samples. Experimental results on the Extended CMU-Seasons and RobotCar-Seasons datasets demonstrate that the proposed method achieves competitive localization performance. On the RobotCar-Seasons dataset, the proposed method achieves the best coarse-level performance under day-night illumination changes. Overall, without relying on depth information, the proposed method achieves localization accuracy comparable to state-of-the-art approaches and surpasses them under certain experimental settings, demonstrating strong cross-domain generalization ability and practical deployment potential.

    中文摘要 I Abstract II 致謝 IV Content V List of Tables VII List of Figures VIII Chapter 1 Introduction 1 1.1 Background 1 1.2 Motivation 4 1.3 Problem 6 1.4 Contribution 8 1.5 Organization 10 Chapter 2 Related Work 11 2.1 Geometry-based Localization 11 2.2 Retrieval-based Localization 14 2.3 Self-supervised Contrastive Learning 17 2.4 Summary 20 Chapter 3 Problem Statement 22 Chapter 4 Methodology 24 4.1 Feature Extraction 25 4.2 Semantic Segmentation Branch 27 4.3 Learnable Semantic-Weighted Contrastive Module 29 4.4 Adversarial Domain Alignment 37 4.5 Overall Training Objective 40 4.6 Image Retrieval Pipeline 42 Chapter 5 Experimental Evaluation 44 5.1 Experimental Data and Settings 44 5.1.1 Datasets 45 5.1.2 Evaluation Metrics 46 5.1.3 Implementation Details 48 5.2 Internal Experiments 49 5.2.1 Effect of Semantic-Guided Masked Pooling Strategies 49 5.2.2 Effect of Contrastive Learning Layers 52 5.2.3 Effect of Retrieval Feature Layers 54 5.2.4 Ablation Study 56 5.3 External Experiments 58 5.3.1 Performance under Various Regional Environments 60 5.3.2 Performance under Various Vegetation Conditions 62 5.3.3 Performance under Various Weather Conditions 64 5.3.4 Performance under Various Illumination Conditions 67 Chapter 6 Conclusion and Future Work 71 References 73

    [1] Relja Arandjelović, Petr Gronat, Akihiko Torii, Tomáš Pajdla, and Josef Sivic, "NetVLAD: CNN Architecture for Weakly Supervised Place Recognition," Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5297-5307, 2016.
    [2] Assia Benbihi, Santiago Arravechia, Matthieu Geist, and Cédric Pradalier, "Image-Based Place Recognition on Bucolic Environment Across Seasons From Semantic Edge Description," Proceedings of the IEEE International Conference on Robotics and Automation, pp. 3032-3038, 2020.
    [3] Hernán Badino, Daniel Huber, and Takeo Kanade, "The CMU Visual Localization Data Set," Carnegie Mellon University, Pittsburgh, PA, USA, 2011. [Online]. Available: http://3dvis.ri.cmu.edu/data-sets/localization
    [4] Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He, "Improved Baselines With Momentum Contrastive Learning," arXiv preprint arXiv:2003.04297, 2020.
    [5] Xinlei Chen, and Kaiming He, "Exploring Simple Siamese Representation Learning," Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15750-15758, 2021.
    [6] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton, "A Simple Framework for Contrastive Learning of Visual Representations," Proceedings of the International Conference on Machine Learning, pp. 1597-1607, 2020.
    [7] Yohann Cabon, Naila Murray, and Martin Humenberger, "Virtual KITTI 2," arXiv preprint arXiv:2001.10773, 2020.
    [8] Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin, "Emerging Properties in Self-Supervised Vision Transformers," Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9650-9660, 2021.
    [9] Xinlei Chen, Saining Xie, and Kaiming He, "An Empirical Study of Training Self-Supervised Vision Transformers," Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9640-9649, 2021.
    [10] Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun, "Vision Meets Robotics: The KITTI Dataset," The International Journal of Robotics Research, vol. 32, no. 11, pp. 1231-1237, 2013.
    [11] Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko, "Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning," Advances in Neural Information Processing Systems, vol. 33, pp. 21271-21284, 2020.
    [12] Fengnian Ge, Yuhui Zhang, Lei Wang, Sonya Coleman, and Dermot Kerr, "Double-Domain Adaptation Semantics for Retrieval-Based Long-Term Visual Localization," IEEE Transactions on Multimedia, vol. 26, pp. 6050-6064, 2024.
    [13] Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick, "Momentum Contrast for Unsupervised Visual Representation Learning," Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9729-9738, 2020.
    [14] Hanjiang Hu, Zhiqiang Qiao, Ming Cheng, Zhen Liu, and Hesheng Wang, "DASGIL: Domain Adaptation for Semantic and Geometric-Aware Image-Based Localization," Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 145-153, 2022.
    [15] Hanjiang Hu, Hesheng Wang, Zhen Liu, Chenguang Yang, Weidong Chen, and Le Xie, "Retrieval-Based Localization Based on Domain-Invariant Feature Learning Under Changing Environments," Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 3684-3689, 2019.
    [16] Hanjiang Hu, Hesheng Wang, Zhen Liu, and Weidong Chen, "Domain-Invariant Similarity Activation Map Contrastive Learning for Retrieval-Based Long-Term Visual Localization," IEEE/CAA Journal of Automatica Sinica, vol. 9, no. 12, pp. 2235-2248, 2022.
    [17] Xudong Mao, Qing Li, Haoran Xie, Raymond Y. K. Lau, Zhen Wang, and Stephen Paul Smolley, "Least Squares Generative Adversarial Networks," Proceedings of the IEEE International Conference on Computer Vision, pp. 2794-2802, 2017.
    [18] William Maddern, Geoffrey Pascoe, Chris Linegar, and Paul Newman, "1 Year, 1000 km: The Oxford RobotCar Dataset," The International Journal of Robotics Research, vol. 36, no. 1, pp. 3-15, 2017.
    [19] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jégou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski, "DINOv2: Learning Robust Visual Features Without Supervision," arXiv preprint arXiv:2304.07193, 2023.
    [20] Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk, "From Coarse to Fine: Robust Hierarchical Localization at Large Scale," Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12716-12725, 2019.
    [21] Torsten Sattler, Will Maddern, Carl Toft, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Tomas Pajdla, Fredrik Kahl, and Torsten Sattler, "Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions," Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8601-8610, 2018.
    [22] Paul-Edouard Sarlin, Ajay Unagar, Måns Larsson, Hugo Germain, Carl Toft, Viktor Larsson, Marc Pollefeys, Vincent Lepetit, Lars Hammarstrand, Fredrik Kahl, and Torsten Sattler, "Back to the Feature: Learning Robust Camera Localization From Pixels to Pose," Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 537-547, 2021.
    [23] Torsten Sattler, Bastian Leibe, and Leif Kobbelt, “Efficient & Effective Prioritized Matching for Large-Scale Image-Based Localization,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 9, pp. 1744–1756, 2017.
    [24] Carl Toft, Will Maddern, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Tomas Pajdla, Fredrik Kahl, and Torsten Sattler, "Long-Term Visual Localization Revisited," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 4, pp. 2074-2088, 2022.
    [25] Shiqi Tang, Yibo Li, Jielin Wan, Yuxuan Li, Bo Zhou, Ronghao Guo, Wei Wang, and Yachun Feng, "TransCNNLoc: End-to-End Pixel-Level Learning for 2D-to-3D Pose Estimation in Dynamic Indoor Scenes," Automation in Construction, vol. 155, p. 104046, 2024.
    [26] Yan Tan, Pei Ji, Yuhui Zhang, Fengnian Ge, and Shihui Zhu, "Learning Robust Representation and Sequence Constraint for Retrieval-Based Long-Term Visual Place Recognition," Pattern Recognition Letters, vol. 172, pp. 56-64, 2024.
    [27] Zhirong Wu, Yuanjun Xiong, Stella X. Yu, and Dahua Lin, "Unsupervised Feature Learning via Non-Parametric Instance Discrimination," Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3733-3742, 2018.
    [28] Jiawei Wang, Yifan Wang, Yuxuan Jin, Chuan Guo, and Xuesong Fan, "ST-PixLoc: A Scene-Agnostic Network for Enhanced Camera Localization," IEEE Robotics and Automation Letters, vol. 8, no. 5, pp. 2785-2792, 2024.
    [29] Zhe Xin, Yinghao Cai, Tao Lu, Xiaoxia Xing, Shaojie Cai, Jixiang Zhang, Yiping Yang, and Yanqing Wang, "Localizing Discriminative Visual Landmarks for Place Recognition," Proceedings of the IEEE International Conference on Robotics and Automation, pp. 5979-5985, 2019.
    [30] Luwei Yang, Zhaopeng Bai, Canglin Tang, Hongdong Li, Yasutaka Furukawa, and Ping Tan, "SANet: Scene Agnostic Network for Camera Localization," Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10979-10988, 2020.

    下載圖示
    校外:立即公開
    QR CODE