| 研究生: |
蘇燿徽 Su, Yao-Hui |
|---|---|
| 論文名稱: |
基於遮罩注意力機制的目標感知孿生網路視覺目標追蹤器 Target-Aware Siamese Networks Based on Masked Attention Mechanism for Visual Object Tracking |
| 指導教授: |
謝明得
Shieh, Ming-Der |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2023 |
| 畢業學年度: | 111 |
| 語文別: | 英文 |
| 論文頁數: | 54 |
| 中文關鍵詞: | 視覺物件追蹤 、孿生神經網路 、注意力機制 、相似度學習 、深度學習 |
| 外文關鍵詞: | visual object tracking, Siamese neural networks, attention mechanism, deep learning, similarity learning |
| 相關次數: | 點閱:158 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
視覺物件追蹤是電腦視覺領域中一個基本且蓬勃發展的任務,視覺物件追蹤的目標是在影片序列中定位和追踪目標物件,並被廣泛應用於多種視覺任務中,如自動駕駛、無人機、人機交互和監視等。然而,當視覺物件追蹤應用於真實世界場景時,由於環境因素的影響,會出現許多挑戰。這些挑戰包括複雜背景、物體形變、物體遮擋等多種不確定因素。因此,克服這些挑戰對於視覺物件追蹤器來說至關重要。
近年來,基於孿生神經網路的視覺物體追蹤演算法在追蹤性能和運算複雜度之間取得了良好的平衡,因此成為了視覺物件追蹤器的熱門網路架構。基於孿生神經網路所構成的物件追蹤器是利用比較欲追蹤目標的模板影像和搜尋區域圖像之間的相似度來實現物件追蹤。然而,在複雜背景的情況下,基於相似度匹配的孿生網路追蹤器可能會因為網路的判別能力不足導致追蹤失敗。本論文在基於孿生神經網路的視覺物件追蹤器中加入了遮罩注意力模組,利用其特性增強目標特徵的表示,進而提升視覺物件追蹤器在複雜背景中的目標區分能力。
最後,本論文使用OTB100做為測試數據集,藉此評估所提出的孿生網路物件追蹤器之性能。實驗結果所示,所提出之遮罩注意力模組可使孿生網路物件追蹤器之性能在複雜背景的情境中有所提升。
Visual object tracking is a popular and fundamental and task in the field of computer vision which aim to locate the target object in a video sequence. Visual object tracking plays a crucial part in various applications, such as unmanned aerial vehicle (UAV), autonomous driving, surveillance, and human-computer interaction, etc.
However, when applying visual object tracking in real-world scenarios, numerous challenges arise due to environmental factors. These challenges include cluttered backgrounds, object deformations, illumination variation, and more. Therefore, addressing these challenges is crucial for modern visual object trackers.
In recent years, visual object tracking algorithms based on Siamese neural networks have achieve a well balance between tracking performance and computational complexity, making it popular frameworks for visual object tracking. Siamese-based tracks rely on comparing similarity between a template image of target and the of the search region image to achieve object tracking. However, these Siamese-based trackers relying on similarity matching may lead to false tracking duo to its insufficient discriminative ability in challenging scenarios such as cluttered backgrounds.
This thesis proposes a Siamese-based tracker with masked attention mechanism which aims to enhance the target features representation and improve the discriminative ability of Siamese-based tracker when dealing with cluttered backgrounds. Finally, we evaluated the performance of the proposed Siamese-based visual object tracker using the OTB100 testing dataset. The experimental results demonstrated improved performance in challenging scenes such cluttered backgrounds scenario.
[1] L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, and P. H. Torr., “Fully-convolutional siamese networks for object tracking”, in ECCV Workshops, 2016.
[2] B. Li, J.J. Yan, W. Wu, Z. Zhu, and X.L. Hu., “High performance visual tracking with siamese region proposal network”, in CVPR, 2018.
[3] B. Li, W. Wu, Q. Wang, F.Y. Zhang, J.L. Xing, and J.J. Yan., “Siamrpn++: Evolution of siamese visual tracking with very deep networks”, in CVPR, 2019.
[4] Z. Zhu, Q. Wang, B. Li, W. Wu, J.J. Yan, and W.M. Hu., “Distractor-aware siamese networks for visual object tracking”, in ECCV, 2018.
[5] G. Wang, C. Luo, X. Sun, Z. Xiong, and W. Zeng., “Tracking by instance detection: A meta-learning approach”, in CVPR, 2020.
[6] H. Lee, S. Choi, and C. Kim., “A memory model based on the Siamese network for long-term tracking”, in ECCV Workshops, 2018.
[7] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin., “Attention is all you need”, in Advances in neural information processing systems, 2017.
[8] D. S. Bolme, J. R. Beveridge, B. A. Draper, and Y. M. Lui., “Visual object tracking using adaptive correlation filters”, in CVPR, 2010.
[9] J. F. Henriques, R. Caseiro, P. Martins, and J. Batista., “High-speed tracking with kernelized correlation filters”, in TPAMI, 2015.
[10] H. Wu, Z. Xu, J. Zhang, W. Yan, and X. Ma., “Face recognition based on convolution siamese networks”, in Int’l Congress on Image and Signal Processing, Bio Medical Engineering and Informatics, 2017.
[11] X. Chen, B. Yan, J. Zhu, D. Wang, X. Yang, and H. Lu., “Transformer tracking”, in CVPR, 2021.
[12] N. Wang, W. Zhou, J. Wang, and H. Li., “Transformer meets tracker: Exploiting temporal context for robust visual tracking”, in CVPR, 2021.
[13] B. Yan, H. Peng, J. Fu, D. Wang, and H. Lu., “Learning spatio-temporal transformer for visual tracking”, in ICCV, 2021.
[14] S. Gao, C. Zhou, and J. Zhang., “Generalized relation modeling for transformer tracking”, in CVPR, 2023.
[15] S. Ren, K. He, R. Girshick, and J. Sun., “Faster r-cnn: Towards real-time object detection with region proposal networks”, in Advances in neural information processing systems, 2015.
[16] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam., “Mobilenets: Efficient convolutional neural networks for mobile vision applications”, arXiv preprint arXiv:1704.04861, 2017.
[17] Q. Wang, L. Zhang, L. Bertinetto, W. Hu, and P. H. Torr., “Fast online object tracking and segmentation: A unifying approach”, in CVPR, 2019.
[18] Z. Chen, B. Zhong, G. Li, S. Zhang, and R. Ji., “Siamese box adaptive network for visual tracking”, in CVPR, 2020.
[19] H. Law and J. Deng., “Cornernet: Detecting objects as paired keypoints”, in ECCV, 2018.
[20] X. Zhou, D. Wang, and P. K., “Objects as points”, arXivpreprintarXiv:1904.07850, 2019.
[21] Z.Tian, C. Shen, H. Chen, and T. He., “FCOS: A simple and strong anchor-free object detector”, in TPAMI, 2022.
[22] D. Guo, J. Wang, Y. Cui, Z. Wang, and S. Chen., “Siamcar: Siamese fully convolutional classification and regression for visual tracking”, in CVPR, 2020.
[23] D. Guo, Y. Shao, Y. Cui, Z. Wang, L. Zhang, and C. Shen., “Graph attention tracking”, in CVPR, 2021.
[24] X. Dong and J. Shen., “Triplet loss in siamese network for object tracking”, in ECCV, 2018.
[25] F. Schroff, D. Kalenichenko, and J. Philibin., “Facenet: A unified embedding for face recognition and clustering”, in CVPR, 2015.
[26] F. Tang, and Q. Ling., “Ranking-based siamese visual tracking”, in CVPR, 2022.
[27] P. Voigtlaender, J. Luiten, P. H. Torr, and B. Leibe., “Siamr-cnn: Visual tracking by re-detection”, in CVPR, 2020.
[28] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly,J. Uszkoreit, and N. Houlsby., An image is worth 16x16 words: Transformers for image recognition at scale”, in ICLR, 2021.
[29] Y. Yu, Y. Xiong, W. Huang, and M. R. Scott., “Deformable siamese attention networks for visual object tracking”, in CVPR, 2020.
[30] P. Blatter, M. Kanakis, M. Danelljan, and L. V. Gool., “Efficient visual tracking with exemplar transformers”, in CVPR, 2023.
[31] A. Krizhevsky, I. Sutskever, and G. E. Hinton., “Imagenet classification with deep convolutional neural networks”, in NIPS, 2012.
[32] J. Fu, J. Liu, H. Tian, Y. Li, Y. Bao, Z. Fang, and H. Lu., “Dual attention network for scene segmentation”, in CVPR, 2019.
[33] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei., “ImageNet large scale visual recognition challenge”, in IJCV, 2015.
[34] T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll´ar, and C. L. Zitnick., “Microsoft coco: Common objects in context”, in ECCV, 2014.
[35] Y. Wu, J. Lim, and M. H. Yang., “Object tracking benchmark”, in TPAMI, 2015.
[36] S. Javed, M. Danelljan, F. S. Khan, M. H. Khan, M. Felsberg, and J. Matas., “Visual object tracking with discriminative filters and siamese networks: A survey and outlook”, arXiv preprint arXiv:2112.02838, 2021.