簡易檢索 / 詳目顯示

研究生: 鄭皓壬
Cheng, Hao-Jen
論文名稱: 基於CARLA模擬器之時序動態感知與速度進度預測閉環自動駕駛
Temporal Motion-Aware Closed-Loop Autonomous Driving with Speed Progress Prediction in CARLA Simulator
指導教授: 楊家輝
Yang, Jar-Ferr
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電腦與通信工程研究所
Institute of Computer & Communication Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 68
中文關鍵詞: 端對端自動駕駛閉環評估時序建模速度進度
外文關鍵詞: autonomous driving, closed-loop evaluation, temporal modeling, speed progress
相關次數: 點閱:21下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 端對端自動駕駛可直接學習由感測器觀測至駕駛決策的映射關係,但在持續變化的交通環境中,如何於僅使用攝影機的閉環系統中有效建模短期時序動態與縱向速度,仍具挑戰性。
    本論文提出一套結合速度進度預測的時序動態感知閉環自動駕駛系統,並於 CARLA 模擬器中進行驗證。所提出的時序動態標記模組擷取相鄰影像幀之間的特徵層級差異,並透過可學習查詢與注意力池化機制,將其壓縮為緊湊的動態標記。此外,一維累積速度進度可將縱向運動與橫向軌跡變化分離,而目標速度與轉向角輔助預測頭則提供控制感知監督,以支援穩定的閉環控制。
    Bench2Drive 實驗結果顯示,相較於復現基線模型,所提出方法在整體閉環駕駛表現以及不同情境下的駕駛能力皆有明顯提升,並能在多種動態交通情境中展現較穩定的路徑規劃與縱向速度調整能力。這些結果驗證了所提出時序動態建模與縱向速度進度預測設計的有效性。

    End-to-end autonomous driving maps sensor observations directly to driving decisions, but modeling short-term temporal dynamics and longitudinal speed remains challenging in camera-only closed-loop systems under continuously changing traffic conditions.
    This thesis proposes a temporal motion-aware closed-loop driving system with speed progress prediction in CARLA. A temporal motion token module captures feature-level differences between adjacent frames and compresses them through learnable queries and attention pooling. One-dimensional cumulative speed progress separates longitudinal motion from lateral trajectory variation, while auxiliary target-speed and steering-angle heads provide control-aware supervision to support stable closed-loop control.
    Experiments on Bench2Drive show that the proposed method achieves better overall driving performance and scenario-level capability than the reproduced baseline. These results demonstrate the effectiveness of the proposed temporal and longitudinal designs.

    摘要 2 Abstract 3 誌謝 4 Contents 5 List of Tables 7 List of Figures 8 Chapter 1 Introduction 9 1.1 Research Background 9 1.2 Motivations 11 1.3 Contributions 12 1.4 Thesis Organization 13 Chapter 2 Related Work 14 2.1 End-to-End Autonomous Driving 14 2.2 Vision-Language Models for Autonomous Driving 15 2.3 Temporal Modeling in Autonomous Driving 19 2.4 Visual Encoders for Driving 20 2.5 Trajectory and Speed Prediction 21 2.6 Closed-Loop Evaluation Benchmarks 23 Chapter 3 The Proposed Closed-Loop Autonomous Driving System 25 3.1 Overview of the Proposed System 26 3.2 Input Representation and Data Pre-processing 28 3.3 Visual Encoder and Query-Based Temporal Motion Token Module 31 3.4 Multi-Task Prediction Head Design 33 3.4.1 Route Prediction Head 34 3.4.2 Speed Progress Prediction Head 34 3.4.3 Target Speed and Steering Angle Heads 36 3.5 Loss Functions 38 3.6 Closed-Loop Inference and Control Module 40 Chapter 4 Experiment Results 43 4.1 Environment Settings and Dataset 43 4.2 Training Details 45 4.3 Evaluation Metrics 46 4.4 Comparison with Existing Methods 49 4.5 Ablation Study 53 4.6 Visualization Results 55 Chapter 5 Conclusions 60 Chapter 6 Future Work 62 References 64

    [1] A. Dosovitskiy, G. Ros, F. Codevilla, A. López, and V. Koltun, “CARLA: An Open Urban Driving Simulator,” in Proc. Conference on Robot Learning (CoRL), vol. 78, 2017.
    [2] X. Jia, Z. Yang, Q. Li, Z. Zhang, and J. Yan, “Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-to-End Autonomous Driving,” in Proc. Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, vol. 37, 2024.
    [3] K. Renz, L. Chen, E. Arani, and O. Sinavski, “SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
    [4] B. Zhang, N. Song, X. Jin, and L. Zhang, “Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
    [5] S. Hu, L. Chen, P. Wu, H. Li, J. Yan, and D. Tao, “ST-P3: End-to-End Vision-Based Autonomous Driving via Spatial-Temporal Feature Learning,” in Proc. European Conference on Computer Vision (ECCV), 2022.
    [6] P. Wu, X. Jia, L. Chen, J. Yan, H. Li, and Y. Qiao, “Trajectory-Guided Control Prediction for End-to-End Autonomous Driving: A Simple Yet Strong Baseline,” in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 35, 2022.
    [7] Y. Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wang, et al., “Planning-Oriented Autonomous Driving,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
    [8] Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Q. Yu, and J. Dai, “BEVFormer: Learning Bird’s-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers,” in Proc. European Conference on Computer Vision (ECCV), 2022.
    [9] X. Jia, P. Wu, L. Chen, J. Xie, C. He, J. Yan, and H. Li, “Think Twice Before Driving: Towards Scalable Decoders for End-to-End Autonomous Driving,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
    [10] X. Jia, Y. Gao, L. Chen, J. Yan, P. L. Liu, and H. Li, “DriveAdapter: Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous Driving,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2023.
    [11] B. Jiang, S. Chen, Q. Xu, B. Liao, J. Chen, H. Zhou, Q. Zhang, W. Liu, C. Huang, and X. Wang, “VAD: Vectorized Scene Representation for Efficient Autonomous Driving,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2023.
    [12] H. Touvron et al., “LLaMA: Open and Efficient Foundation Language Models,” arXiv preprint arXiv:2302.13971, 2023.
    [13] H. Touvron et al., “Llama 2: Open Foundation and Fine-Tuned Chat Models,” arXiv preprint arXiv:2307.09288, 2023.
    [14] Z. Song, C. Jia, L. Liu, H. Pan, Y. Zhang, J. Wang, X. Zhang, S. Xu, L. Yang, and Y. Luo, “Don’t Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Driving,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
    [15] S. Shang, Y. Chen, Y. Wang, Y. Li, and Z. Zhang, “DriveDPO: Policy Learning via Safety DPO for End-to-End Autonomous Driving,” in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 38, 2025.
    [16] X. Jia, J. You, Z. Zhang, and J. Yan, “DriveTransformer: Unified Transformer for Scalable End-to-End Autonomous Driving,” in Proc. International Conference on Learning Representations (ICLR), 2025.
    [17] T. Wang, C. Zhang, X. Qu, K. Li, W. Liu, and C. Huang, “DiffAD: A Unified Diffusion Modeling Approach for Autonomous Driving,” arXiv preprint arXiv:2503.12170, 2025.
    [18] Z. Yang, X. Jia, Q. Li, X. Yang, M. Yao, and J. Yan, “Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2),” in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 38, 2025.
    [19] Z. Yang, Y. Chai, X. Jia, Q. Li, Y. Shao, X. Zhu, H. Su, and J. Yan, “DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026.
    [20] Y. Li, K. Xiong, X. Guo, F. Li, S. Yan, G. Xu, L. Zhou, L. Chen, H. Sun, B. Wang, K. Ma, G. Chen, H. Ye, W. Liu, and X. Wang, “ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving,” in Proc. International Conference on Learning Representations (ICLR), 2026.
    [21] H. Fu, D. Zhang, Z. Zhao, J. Cui, D. Liang, C. Zhang, D. Zhang, H. Xie, B. Wang, and X. Bai, “ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2025.
    [22] H. Fu, D. Zhang, Z. Zhao, J. Cui, H. Xie, B. Wang, G. Chen, D. Liang, and X. Bai, “MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning,” arXiv preprint arXiv:2512.13636, 2025.
    [23] Z. Xiong, X. Ye, B. Yaman, S. Cheng, Y. Lu, J. Luo, N. Jacobs, and L. Ren, “UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving,” to appear in Proc. European Conference on Computer Vision (ECCV), 2026.
    [24] J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models,” in Proc. International Conference on Machine Learning (ICML), vol. 202, 2023.
    [25] J. Beißwenger, “PDM-Lite: A Rule-Based Planner for CARLA Leaderboard 2.0,” University of Tübingen, Tech. Rep., 2024.
    [26] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-Rank Adaptation of Large Language Models,” in Proc. International Conference on Learning Representations (ICLR), 2022.
    [27] I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” in Proc. International Conference on Learning Representations (ICLR), 2019.
    [28] H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee, “LLaVA-NeXT: Improved Reasoning, OCR, and World Knowledge,” LLaVA Project Blog, 2024.
    [29] M. Oquab et al., “DINOv2: Learning Robust Visual Features without Supervision,” Transactions on Machine Learning Research, 2024.
    [30] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017.
    [31] L. N. Smith and N. Topin, “Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates,” in Proc. SPIE Artificial Intelligence and Machine Learning for Multi-Domain Operations Applications, vol. 11006, Art. no. 1100612, 2019, doi: 10.1117/12.2520589.
    [32] K. Chitta, A. Prakash, B. Jaeger, Z. Yu, K. Renz, and A. Geiger, “TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 11, 2023.
    [33] B. Jaeger, K. Chitta, and A. Geiger, “Hidden Biases of End-to-End Driving Models,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2023.
    [34] C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, J. Beißwenger, P. Luo, A. Geiger, and H. Li, “DriveLM: Driving with Graph Visual Question Answering,” in Proc. European Conference on Computer Vision (ECCV), 2024.
    [35] Z. Xu, Y. Bai, Y. Zhang, Z. Li, F. Xia, K.-Y. K. Wong, J. Wang, and H. Zhao, “DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving,” in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
    [36] Q. Li, X. Jia, S. Wang, and J. Yan, “Think2Drive: Efficient Reinforcement Learning by Thinking with Latent World Model for Autonomous Driving (in CARLA-v2),” in Proc. European Conference on Computer Vision (ECCV), 2024.

    下載圖示
    校外:立即公開
    QR CODE