簡易檢索 / 詳目顯示

研究生: 黃彥淇
Huang, Yen-Chi
論文名稱: 結合學習式自適應雜訊之強化擴散模型的單目視訊深度估計
Diffusion-Based Monocular Video Depth Estimation Enhanced with Learned Adaptive Noise
指導教授: 楊家輝
Yang, Jar-Ferr
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電腦與通信工程研究所
Institute of Computer & Communication Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 72
中文關鍵詞: 單目深度估計 、影片深度估計 、擴散模型 、自適應雜訊 、時間一致性
外文關鍵詞: monocular depth estimation, video depth estimation, diffusion models, adaptive noise, temporal consistency
相關次數: 點閱:108  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 單目深度估計在自動駕駛、機器人導航以及三維場景理解等應用中受到廣泛關注。相較於雙目或基於光達的方法,單目深度估計僅依賴單一彩色相機,因此具有更高的成本效益與更容易部署的優勢。然而,由於深度本身存在固有歧義,該問題仍屬於典型的不適定問題,且在影片序列中維持時間一致性仍然具有挑戰性。
    近年來,擴散模型在單目深度估計任務中展現出優異的性能。然而,目前多數方法(例如 Marigold法)採用固定的噪聲排程策略,缺乏對空間結構與時間變化的自適應能力,因而在影片應用中容易產生邊界偽影與時間閃爍現象。為了解決此問題,本論文提出一種基於擴散模型的單目深度估計框架,並引入學習式自適應噪聲機制以適用於影片序列。所提出的方法可根據影像的空間特徵與擴散時間步動態調整噪聲排程,並結合時間深度潛變量傳播機制,在不使用光流或循環神經網路的情況下提升跨幀一致性。
    實驗結果顯示,所提出的方法在深度品質與時間一致性方面均穩定優於Marigold 法。相較於Depth Anything 法,本方法在較小的訓練規模下仍展現具競爭力的表現,並在時間一致性上具有明顯優勢,特別適用於基於影片的深度估計任務。

    Monocular depth estimation is widely used in autonomous driving, robotic navigation, and 3D scene understanding. However, it remains an ill-posed problem due to depth ambiguity, and achieving temporal consistency in video sequences is still challenging.
    Recently, diffusion models have shown promising results in monocular depth estimation. However, most methods, such as Marigold, rely on fixed noise schedules that cannot adapt to spatial structures or temporal variations, often causing boundary artifacts and temporal flickering. To address this limitation, this thesis proposes a diffusion-based monocular depth estimation framework with a learned adaptive noise mechanism and temporal depth latent propagation, improving cross-frame consistency without optical flow or recurrent networks.
    Experimental results show that the proposed method outperforms Marigold in both depth quality and temporal consistency. It also achieves competitive performance with Depth Anything using a much smaller training setup while providing superior temporal consistency.

    摘要 II Abstract III 誌謝 IV Contents V List of Tables VII List of Figures VIII Chapter 1 Introduction 1 1.1 Research Background 2 1.2 Motivations 3 1.3 Thesis Organization 4 Chapter 2 Related Work 5 2.1 Conventional Depth Estimation Methods 6 2.2 Monocular Depth Estimation in the Deep Learning Era 7 2.3 Video Depth Estimation and Temporal Consistency 9 2.4 Diffusion Models 11 2.5 Diffusion-Based Monocular Depth Estimation 13 2.6 Adaptive Noise Strategies in Diffusion Models 15 Chapter 3 The Proposed Video Monocular Depth Estimation System 18 3.1 Overview of the Proposed Depth Estimation Framework 19 3.2 Baseline Diffusion Model 20 3.2.1 Diffusion Training Objective 21 3.2.2 Latent Encoder and Decoder 22 3.2.3 Latent Denoising U-Net 24 3.2.4 LoRA Fine-Tuning 26 3.2.5 Affine-Invariant Depth Representation 27 3.3 Temporal Depth Latent Concatenation 28 3.4 Learnable Noise Generator for Depth Diffusion 30 3.4.1 Image-Guided Noise Conditioning Encoder 31 3.4.2 Learned Noise Modulation Function 33 3.4.3 Adaptive Noise Generation Formulation 36 3.5 Inference Procedure 37 Chapter 4 Experiment Results 40 4.1 Environment Setup and Datasets 40 4.2 Training Details and Hyperparameter Settings 42 4.3 Evaluation Metrics 43 4.4 Comparison with Other Methods 45 4.5 Ablation Study 49 4.6 Visualization of Prediction Results 51 Chapter 5 Conclusions 58 Chapter 6 Future Work 59 References 60

    [1] David Eigen, C. Puhrsch, and R. Fergus, “Depth Map Prediction from a Single Image using a Multi-Scale Deep Network,” in Advances in Neural Information Processing Systems (NeurIPS), 2014.
    [2] Ren Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun, “Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 3, pp. 1623–1637, 2022.
    [3] Bingxin Ke, A. Obukhov, S. Huang, N. Metzger, R. C. Daudt, and K. Schindler, “Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation,” in CVPR, 2024.
    [4] Jiaming Xu, Z. Duan, N. Wang, Y. Zhang, and H. Shen, “DiffusionDepth: Diffusion Denoising Approach for Monocular Depth Estimation,” in AAAI Conference on Artificial Intelligence, 2023.
    [5] R. Szeliski, Computer Vision: Algorithms and Applications, 2nd ed. Cham, Switzerland: Springer, 2022.
    [6] Andreas Geiger, P. Lenz, and R. Urtasun, “Are We Ready for Autonomous Driving? The KITTI Vision Benchmark Suite,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2012.
    [7] Richard Hartley and A. Zisserman, Multiple View Geometry in Computer Vision, 2nd ed. Cambridge, U.K.: Cambridge University Press, 2004.
    [8] Johannes L. Schönberger and J.-M. Frahm, “Structure-from-Motion Revisited,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2016.
    [9] Jin Han Lee, M. Han, D. Ko, and I. S. Kweon, “From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2019.
    [10] Shariq Farooq Bhat, I. Alhashim, and P. Wonka, “AdaBins: Depth Estimation Using Adaptive Bins,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2021.
    [11] Ren Ranftl, A. Bochkovskiy, and V. Koltun, “Vision Transformers for Dense Prediction,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
    [12] H. Zhang, C. Shen, Y. Li, Y. Cao, Y. Liu, and Y. Yan, "Exploiting Temporal Consistency for Real-Time Video Depth Estimation,"arXiv:1908.03706, 2019.
    [13] J. Shao, Yuanbo Yang, Hongyu Zhou, Youmin Zhang, Yujun Shen, Vitor Guizilini, Yue Wang, Matteo Poggi, Yiyi Liao, "Learning Temporally Consistent Video Depth from Video Diffusion Priors, "arXiv:2406.01493, 2024.
    [14] J. J. Leonard and H. F. Durrant-Whyte, “Directed Sonar Sensing for Mobile Robot Navigation,” Kluwer Academic Publishers, 1992.
    [15] R. Lange and P. Seitz, “Solid-State Time-of-Flight Range Camera,” IEEE Journal of Quantum Electronics, vol. 37, no. 3, pp. 390–397, 2001.
    [16] D. Scharstein and R. Szeliski, “A Taxonomy and Evaluation of Dense Two-Frame Stereo Correspondence Algorithms,” International Journal of Computer Vision, vol. 47, no. 1–3, pp. 7–42, 2002.
    [17] C. Tomasi and T. Kanade, “Shape and Motion from Image Streams under Orthography: A Factorization Method,” International Journal of Computer Vision, vol. 9, no. 2, pp. 137–154, 1992.
    [18] S. M. Seitz, B. Curless, J. Diebel, D. Scharstein, and R. Szeliski, “A Comparison andEvaluation of Multi-View Stereo Reconstruction Algorithms,” Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2006.
    [19] A. Saxena, S. H. Chung, and A. Y. Ng, “Learning Depth from Single Monocular Images,” Advances in Neural Information Processing Systems, vol. 18, 2005.
    [20] A. Saxena, J. Schulte, and A. Y. Ng, “Depth Estimation Using Monocular and Stereo Cues,” IJCAI, vol. 7, pp. 2197–2203, 2007.
    [21] C. Godard, O. Mac Aodha, M. Firman, and G. Brostow, “Digging Into Self Supervised Monocular Depth Estimation,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2019.
    [22] A. Gordon, H. Li, R. Jonschkowski, and A. Angelova, “Depth from Videos in the Wild: Unsupervised Monocular Depth Learning from Unknown Cameras,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2019.
    [23] S. Kumar, Y. Dai, and H. Li, “Monocular Dense 3D Reconstruction of a Complex Dynamic Scene from Two Perspective Frames,” in Proc. IEEE International Conference on Computer Vision (ICCV), 2017.
    [24] Y. Zhang, M. Tang, Z. Ding, and J. Fu, “TC-Depth: Temporally Consistent Depth Prediction for Video Sequences,” in Pattern Recognition, vol. 123, 2022.
    [25] J. Ho, A. Jain, and P. Abbeel, “Denoising Diffusion Probabilistic Models,” in Advances in Neural Information Processing Systems (NeurIPS), 2020.
    [26] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-Resolution Image Synthesis with Latent Diffusion Models,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
    [27] X. Xu, Y. Yin, M. Xu, C. Shen, and Y. Wang, “GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image,” in Advances in Neural Information Processing Systems (NeurIPS), 2024.
    [28] A. Nichol and P. Dhariwal, “Improved Denoising Diffusion Probabilistic Models,” in Proc. International Conference on Machine Learning (ICML), 2021.
    [29] M. Lee, H. Kim, and S. Kim, “MuLAN: Multi Loss Adaptation with Learnable Parameter-Based Noise Scheduling for Diffusion Models,” in Proc. IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2023.
    [30] T. Chen, R. Zhang, and G. Hinton, “Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning,” in International Conference on Learning Representations (ICLR), 2023.
    [31] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen,“LoRA: Low-Rank Adaptation of Large Language Models,” in International Conference on Learning Representations (ICLR), 2022.
    [32] G. Gaidon, Q. Wang, Y. Cabon, and E. Vig, "Virtual Worlds as Proxy for Multi-Object Tracking Analysis," in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
    [33] M. Li, A. Dosovitskiy, D. Koltun, and T. Funkhouser, "InteriorNet: Mega-scale Multi-sensor Photo-realistic Indoor Scenes Dataset," in British Machine Vision Conference (BMVC), 2018.
    [34] N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, "Indoor Segmentation and Support Inference from RGBD Images," in Proc. European Conference on Computer Vision (ECCV), 2012.
    [35] T. Koch, L. Liebel, M. Körner, and F. Fraundorfer, “Comparison of monocular depth estimation methods using geometrically relevant metrics on the IBims-1 dataset,” Computer Vision and Image Understanding, vol. 191, p. 102877, 2020, doi: 10.1016/j.cviu.2019.102877.
    [36] S. Song, S. P. Lichtenberg, and J. Xiao, "SUN RGB-D: A RGB-D Scene Understanding Benchmark Suite," in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
    [37] G. Ros, L. Sellart, J. Materzynska, D. Vazquez, and A. M. Lopez, "The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes," in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
    [38] J. He, H. Li, W. Yin, Y. Liang, L. Li, K. Zhou, H. Zhang, B. Liu, and Y.-C. Chen, "Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction," in International Conference on Learning Representations (ICLR), 2025.
    [39] L. Yang, Z. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao, "Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data," in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.

    下載圖示
    校外:立即公開
    QR CODE