| 研究生: |
陳竑曄 Chen, Hung-Yeh |
|---|---|
| 論文名稱: |
重新檢視SORT系列多物件追蹤器在斜俯角監視下的長寬比恆定假設 Revisiting the Constant Aspect-Ratio Assumption in SORT-Family Multi-Object Tracking under Oblique-Angle Surveillance |
| 指導教授: |
賀保羅
Horton, Paul |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 醫學資訊研究所 Institute of Medical Informatics |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 72 |
| 中文關鍵詞: | 多目標追蹤 、ByteTrack 、卡爾曼濾波器 、長寬比動態 、狀態表示法 、交通監控 、CCTV |
| 外文關鍵詞: | multi-object tracking, ByteTrack, Kalman filter, aspect-ratio dynamics, state representation, traffic surveillance, CCTV |
| 相關次數: | 點閱:113 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
許多 SORT 系列多物件追蹤器使用卡爾曼濾波器預測邊界框,並假設邊界框的長寬比不變或只會小幅變化。當物體形狀穩定時,這項假設有助於抑制偵測框的抖動;但在交通監視影像中,畫面上實際觀察到的車輛邊界框可能因進出畫面而被裁切、因轉彎而改變方向,或在斜俯角視角下隨位置產生透視變化。本論文探討這項假設與實際邊界框變化不一致時,會如何影響追蹤結果,以及解除長寬比限制是否足以改善問題。
本研究以 ByteTrack 為實驗平台,在固定偵測器、資料關聯方法與追蹤參數的情況下,比較原本的 XYAH 卡爾曼濾波器與不限制長寬比的 XYWH 表示法。真值統計、控制模擬及追蹤器內部預測結果顯示,當長寬比受到限制時,預測框的形狀會跟不上實際邊界框的變化,因而降低兩者的交並比(IoU)。這項誤差主要透過兩種方式影響追蹤:輸出框在較嚴格的定位門檻下無法正確匹配,以及預測框的配對餘裕縮小,造成較多軌跡中斷。身分切換的發生則比原先預期少,並不是主要影響。
將 XYAH 改為 XYWH 後,自建台灣 CCTV 驗證集的 HOTA 由 0.702 提升至 0.715,UA-DETRAC 40 段測試序列的 HOTA 由 0.615 提升至 0.638。在 UA-DETRAC 上,軌跡破碎數由 1,454 降至 1,155,且 40 段序列的 HOTA 均有改善。另一組對照實驗只改變 XYAH 中的 ar_noise_scale 設定參數,結果便與 XYWH 幾乎相同,表示改善主要來自解除長寬比限制,而不是更換追蹤器或資料關聯方法。相較之下,長寬比近乎不變的行人類別幾乎沒有追蹤改善。本研究的逐序列統計證據主要來自 UA-DETRAC 的 40 段測試序列;僅含 7 段序列的台灣驗證集主要提供目標應用場域的實證,以及相同場景中的行人對照,而不作為獨立的顯著性證據。
XYWH 是既有的邊界框表示法,並非本論文提出的新追蹤器。本研究的主要貢獻是找出長寬比限制失效的機制,並透過控制比較與 XYAH-loose 對照實驗,確認主要改善來自放寬長寬比限制,進一步說明這項修改適用與不適用的情況。斜俯角監視影像中的邊界框形狀變化較為頻繁,但問題的根源是長寬比本身的變化,而不只是相機角度。
Many SORT-family multi-object trackers predict bounding boxes with a Kalman filter that assumes the aspect ratio changes little or not at all. This assumption is reasonable when object shape is stable, but it does not always match traffic-surveillance video. The visible bounding box of a vehicle can change shape when the vehicle is clipped by the frame boundary, turns, or undergoes perspective change while moving through an oblique camera view. This thesis studies how this mismatch affects tracking and whether removing the aspect-ratio constraint is sufficient to address it.
We compare ByteTrack's standard XYAH Kalman filter with the aspect-ratio-free XYWH filter while holding the detector, association method, and tracking parameters fixed. Ground-truth statistics, controlled simulations, and instrumented tracker predictions show that a constrained aspect ratio causes the predicted box to lag behind the observed shape. The resulting IoU loss affects tracking in two main ways: inaccurate output boxes fail stricter localization thresholds, and inaccurate predictions reduce the matching margin and increase track fragmentation. Identity switches are a much less common consequence than initially expected.
Replacing XYAH with XYWH improves HOTA from 0.702 to 0.715 on our Taiwan CCTV validation set and from 0.615 to 0.638 on the 40-sequence UA-DETRAC test set. On UA-DETRAC, fragmentation decreases from 1,454 to 1,155, and every test sequence improves in HOTA. A control that changes only the ar_noise_scale configuration parameter in XYAH performs almost identically to XYWH, showing that the gain comes mainly from releasing the aspect-ratio constraint. The pedestrian control class, whose aspect ratio is nearly constant, shows almost no tracking improvement. The main sequence-level statistical evidence comes from the 40-sequence UA-DETRAC test set; the seven-sequence Taiwan validation set instead provides evidence from the target application setting and a within-scene pedestrian comparison.
XYWH is an existing box-state representation; this thesis does not propose it as a new tracking method. The contribution is to identify the failure mechanism, isolate its main source, and clarify when the representation change is useful. Oblique-angle surveillance is a setting in which bounding-box shape changes are frequent, but the failure itself is driven by aspect-ratio dynamics rather than camera angle alone.
[1] N. Aharon, R. Orfaig, and B.-Z. Bobrovsky, “BoT-SORT: Robust associations multi-pedestrian tracking,” arXiv preprint arXiv:2206.14651, 2022.
[2] K. Bernardin and R. Stiefelhagen, “Evaluating multiple object tracking performance: The CLEAR MOT metrics,” EURASIP Journal on Image and Video Processing, vol. 2008, pp. 1–10, 2008.
[3] A. Bewley, Z. Ge, L. Ott, F. Ramos, and B. Upcroft, “Simple online and realtime tracking,” IEEE International Conference on Image Processing (ICIP), 2016.
[4] M. Broström, BoxMOT: Pluggable multi-object tracking library, 2024. [Online]. Available: https://github.com/mikel-brostrom/boxmot
[5] J. Cao, J. Pang, X. Weng, R. Khirodkar, and K. Kitani, “Observation-centric SORT: Rethinking SORT for robust multi-object tracking,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
[6] D. Du, Y. Qi, H. Yu, Y. Yang, K. Duan, G. Li, W. Zhang, Q. Huang, and Q. Tian, “The unmanned aerial vehicle benchmark: Object detection and tracking,” European Conference on Computer Vision (ECCV), 2018, pp. 370–386.
[7] Y. Du, Z. Zhao, Y. Song, Y. Zhao, F. Su, T. Gong, and H. Meng, “StrongSORT: Make DeepSORT great again,” IEEE Transactions on Multimedia, 2023.
[8] G. D. Evangelidis and E. Z. Psarakis, “Parametric image alignment using enhanced correlation coefficient maximization,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 30, no. 10, pp. 1858–1865, 2008.
[9] G. Jocher and J. Qiu, “Ultralytics YOLO11,” 2024. [Online]. Available: https://docs.ultralytics.com/models/yolo11/
[10] R. E. Kalman, “A new approach to linear filtering and prediction problems,” 1960, vol. 82, pp. 35–45.
[11] Y. Lin, S. Lockyer, M. Sui, L. Gan, F. Stanek, M. Zarbock, W. Li, A. Evans, and N. Zhang, “RoundaboutHD: High-resolution real-world urban environment benchmark for multi-camera vehicle tracking,” arXiv preprint arXiv:2507.08729, 2025.
[12] L. Liu, Y. Cheng, Z. Deng, S. Wang, D. Chen, X. Hu, P. Liò, C.-B. Schönlieb, and A. Aviles-Rivero, “TrafficMOT: A challenging dataset for multi-object tracking in complex traffic scenarios,” arXiv preprint arXiv:2311.18839, 2023.
[13] J. Luiten and A. Osěp, TrackEval: Multi-object tracking evaluation, 2021. [Online]. Available: https://github.com/JonathonLuiten/TrackEval
[14] J. Luiten, A. Osěp, P. Dendorfer, P. Torr, A. Geiger, and L. Leal-Taixé, “HOTA: A higher order metric for evaluating multi-object tracking,” International Journal of Computer Vision, vol. 129, pp. 548–578, 2021.
[15] G. Maggiolino, A. Ahmad, J. Cao, and K. Kitani, “Deep OC-SORT: Multi-pedestrian tracking by adaptive re-identification,” IEEE International Conference on Image Processing (ICIP), 2023.
[16] A. Milan, L. Leal-Taixé, I. Reid, S. Roth, and K. Schindler, “MOT16: A benchmark for multi-object tracking,” arXiv preprint arXiv:1603.00831, 2016.
[17] Ministry of Transportation and Communications, Taiwan, TDX: Transport data eXchange platform, https://tdx.transportdata.tw/, Accessed: 2026-06-21, 2026.
[18] T. T. T. Nguyen, V. L. Ho, T. T. A. Nguyen, T. B. Phung, and T. D. Nguyen, “An integrated approach for multi-object detection and tracking in traffic monitoring using YOLOv9c and ByteTrack,” Journal of Technical Education Science, vol. 21, no. 1, pp. 81–90, 2026. DOI: 10.54644/jte.2026.2077
[19] E. Ristani, F. Solera, R. Zou, R. Cucchiara, and C. Tomasi, “Performance measures and a data set for multi-target, multi-camera tracking,” European Conference on Computer Vision (ECCV) Workshops, 2016.
[20] Z. Tang, M. Naphade, M.-Y. Liu, X. Yang, S. Birchfield, S. Wang, R. Kumar, D. Anastasiu, and J.-N. Hwang, “CityFlow: A city-scale benchmark for multi-target multi-camera vehicle tracking and re-identification,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 8797–8806.
[21] Z. Wang, L. Zheng, Y. Liu, Y. Li, and S. Wang, “Towards real-time multi-object tracking,” European Conference on Computer Vision (ECCV), 2020.
[22] L. Wen, D. Du, Z. Cai, Z. Lei, M.-C. Chang, H. Qi, J. Lim, M.-H. Yang, and S. Lyu, “UA-DETRAC: A new benchmark and protocol for multi-object detection and tracking,” Computer Vision and Image Understanding, vol. 193, 2020.
[23] N. Wojke, A. Bewley, and D. Paulus, “Simple online and realtime tracking with a deep association metric,” arXiv preprint arXiv:1703.07402, 2017.
[24] G.-S. Xia, X. Bai, J. Ding, Z. Zhu, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang, “DOTA: A large-scale dataset for object detection in aerial images,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 3974–3983.
[25] L. Yang and T. Hong, “Real-time detection of near-miss events and risk assessment in urban traffic using multi-object tracking and bird’s eye view mapping,” Future Transportation, vol. 6, no. 2, p. 80, 2026. DOI: 10.3390/futuretransp6020080
[26] M. Yang, G. Han, B. Yan, W. Zhang, J. Qi, H. Lu, and D. Wang, “Hybrid-SORT: Weak cues matter for online multi-object tracking,” AAAI Conference on Artificial Intelligence (AAAI), 2024.
[27] F. Zeng, B. Dong, Y. Zhang, T. Wang, X. Zhang, and Y. Wei, “MOTR: End-to-end multiple-object tracking with transformer,” European Conference on Computer Vision (ECCV), 2022.
[28] Y. Zhang, P. Sun, Y. Jiang, D. Yu, F. Weng, Z. Yuan, P. Luo, W. Liu, and X. Wang, “ByteTrack: Multi-object tracking by associating every detection box,” European Conference on Computer Vision (ECCV), 2022.
[29] Y. Zhang, C. Wang, X. Wang, W. Zeng, and W. Liu, “FairMOT: On the fairness of detection and re-identification in multiple object tracking,” International Journal of Computer Vision, vol. 129, no. 11, pp. 3069–3087, 2021.
[30] M. Zhu, S. Zhang, Y. Zhong, P. Lu, H. Peng, and J. Lenneman, “Monocular 3D vehicle detection using uncalibrated traffic cameras through homography,” IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021.
[31] P. Zhu, L. Wen, D. Du, X. Bian, H. Fan, Q. Hu, and H. Ling, “Detection and tracking meet drones challenge,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 44, no. 11, pp. 7380–7399, 2021. DOI: 10.1109/TPAMI.2021.3119563