| 研究生: |
陳貞霓 Chen, Chen-Ni |
|---|---|
| 論文名稱: |
球如何飛行與為何飛行? 排球軌跡切割與分類 How It Flies and Why It Flies? Volleyball Trajectory Segmentation and Classification |
| 指導教授: |
朱威達
Chu, Wei-Ta |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 人工智慧科技碩士學位學程 Graduate Program of Artificial Intelligence |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 英文 |
| 論文頁數: | 30 |
| 中文關鍵詞: | 排球影片 、軌跡切割 、軌跡分類 |
| 外文關鍵詞: | volleyball video, trajectory segmentation, trajectory classification |
| 相關次數: | 點閱:269 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
在排球比賽中,將包含一長串多次來回的排球軌跡進行適當地分割與分類後,這些軌跡片段將有助於更進階的排球比賽分析。過去排球比賽影片的研究大部分將軌跡分割以及軌跡分類當成兩個不同的問題在處理。基於三維軌跡的分析更是相較少數。在本論文中,我們專注於利用球體三維運動軌跡來協助排球影片分析。我們同時考慮軌跡分割與軌跡分類,視為單一問題處理。我們從兩個不同的視角偵測球體的位置,接著建構出球體的三維軌跡。然後結合球體的三維位置 (x,y,z) 以及移動方向資訊當做輸入,提出了基於BERT (Bidirectional Encoder Representation for Transformer) 的軌跡切割與分類方法。我們根據國際排球協會的排球資訊系統,定義六種軌跡類別:發球、接發球、舉球、攻擊、守備、攔網。將一長串連續軌跡中的每一幀分成這六類中的其中一類。如此,包含多次來回的一長串軌跡將會被相同類別的連續幀適當的分割成有意義的片段。我們相信這是第一個利用語言模型技術來分析排球軌跡的研究。軌跡的切割與分類結果,讓我們可以進行更進階的排球比賽分析。
In a volleyball game, how the ball moves is very important in tactics design and performance evaluation. After appropriate segmenting and classifying a long ball trajectory which showing back and forth flying, the trajectory segments can enable more advanced volleyball analysis. Most previous works separated trajectory segmentation and classification into two problems, and trajectory classification was usually based on ad hoc rules. In this thesis, we focus on how ball trajectories can benefit volleyball video analysis and jointly formulate segmentation and classification as a single problem. Based on videos captured by two cameras from different viewpoints, we detect the volleyball and construct 3D ball trajectories. We then jointly consider 3D position information (x,y,z) and movement information as inputs, and propose a trajectory segmentation and classification method based on BERT (Bidirectional Encoder Representation for Transformer). For the definitions of trajectory segment classes, we refer to the Volleyball Information System (VIS) of the International Volleyball Federation (FIVB). The VIS is used to grade each player's skills based on 6 categories (serve, receive, set, spike, block, and dig). The volleyball at each frame can be categorized into one of six trajectory classes, and a long ball trajectory showing the ball being hit back and forth is appropriately segmented. We believe that this is a very first study adopting the language model technique to analyze ball trajectories, and results of trajectory segmentation and classification can enable more advanced volleyball analysis.
[1] Luai Al Shalabi, Zyad Shaaban, and Basel Kasasbeh. Data mining: A preprocessing engine. Journal of Computer Science, 2(9):735–739, 2006.
[2] Timur Bagautdinov, Alexandre Alahi, François Fleuret, Pascal Fua, and Silvio Savarese. Social scene understanding: End-to-end multi-person action localization and collective activity recognition. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 4315–4324, 2017.
[3] Alexey Bochkovskiy, Chien-Yao Wang, and Hong-Yuan Mark Liao. Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934, 2020.
[4] Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pages 6299–6308, 2017.
[5] Bodhisattwa Chakraborty and Sukadev Meher. A trajectory-based ball detection and tracking system with applications to shot-type identification in volleyball videos. In Proceedings of International Conference on Signal Processing and Communications, pages 1–5, 2012.
[6] Hua-Tsung Chen, Hsuan-Shen Chen, and Suh-Yin Lee. Physics-based ball tracking in volleyball videos with its applications to set type recognition and action detection. In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, volume 1, pages I–1097. IEEE, 2007.
[7] Hua-Tsung Chen, Wen-Jiin Tsai, Suh-Yin Lee, and Jen-Yu Yu. Ball tracking and 3d trajectory approximation with applications to tactics analysis from single-camera volleyball sequences. Multimedia Tools and Applications, 60(3):641–667, 2012.
[8] Junwen Chen, Haiting Hao, Hanbin Hong, and Yu Kong. Rit-18: A novel dataset for compositional group activity understanding. In Proceedings of IEEE International Conference on Computer Vision and Pattern Recognition Workshops, 2020.
[9] Xina Cheng, Norikazu Ikoma, Masaaki Honda, and Takeshi Ikenaga. Multi-view 3d ball tracking with abrupt motion adaptive system model, anti-occlusion observation and spatial density based recovery in sports analysis. IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences, 100(5):1215–1225, 2017.
[10] Xina Cheng, Norikazu Ikoma, Masaaki Honda, and Takeshi Ikenaga. Simultaneous physical and conceptual ball state estimation in volleyball game analysis. In Proceedings of IEEE International Conference on Visual Communications and Image Process- ing, 2017.
[11] Wei-Ta Chu and Wen-Ho Tsai. Modeling spatiotemporal relationships between moving objects for event tactics analysis in tennis videos. Multimedia Tools and Applications, 50(1):149–171, 2010.
[12] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina N. Toutanova. Bert: Pre- training of deep bidirectional transformers for language understanding. In Proceedings of The North American Chapter of the Association for Computational Linguistics, 2019.
[13] Dirk Farin, Susanne Krabbe, Wolfgang Effelsberg, and Peter H.N. de With. Robust camera calibration for sport videos using court models. In Proceedings of SPIE Storage and Retrieval Methods and Applications for Multimedia, 2004.
[14] Kirill Gavrilyuk, Ryan Sanford, Mehrsan Javan, and Cees Snoek. Actor-transformers for group activity recognition. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 839–848, 2020.
[15] Richard Hartley and Andrew Zisserman. Multiple View Geometry in Computer Vision. Cambridge University Press, 2003.
[16] Yu-Chuan Huang, I-No Liao, Ching-Hsuan Chen, Tsì-Uí İk, and Wen-Chih Peng. Tracknet: a deep learning network for tracking high-speed and tiny objects in sports applications. In Proceedings of IEEE International Conference on Advanced Video and Signal Based Surveillance, pages 1–8. IEEE, 2019.
[17] Mostafa S. Ibrahim, Srikanth Muralidharan, Zhiwei Deng, Arash Vahdat, and Greg Mori. A hierarchical deep temporal model for group activity recognition. In Proceedings of IEEE International Conference on Computer Vision and Pattern Recognition, 2016.
[18] Dong-Hyun Lee. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Proceedings of ICML Workshop on Challenges in Representation Learning, 2013.
[19] Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision, pages 2980–2988, 2017.
[20] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ra- manan, Piotr Dollar, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In Proceedings of European Conference on Computer Vision, 2014.
[21] Jun Liu, Amir Shahroudy, Dong Xu, Alex C. Kot, and Gang Wang. Skeleton-based action recognition using spatio-temporal lstm network with trust gates. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(12):3007–3021, 2017.
[22] Yang Liu, Shuyi Huang, Xina Cheng, and Takeshi Ikenaga. 3d global trajectory and multi-view local motion combined player action recognition in volleyball analysis. In Proceedings of Pacific Rim Conference on Multimedia, pages 134–144. Springer, 2018.
[23] Antoine Miech, Ivan Laptev, and Josef Sivic. Learnable pooling with context gating for video classification. arXiv preprint arXiv:1706.06905, 2017.
[24] Rizard Renanda Adhi Pramono, Yie Tarng Chen, and Wen Hsien Fang. Empowering relational network by self-attention augmented conditional random fields for group activity recognition. In Proceedings of European Conference on Computer Vision, pages 71–90. Springer, 2020.
[25] Ryan Sanford, Siavash Gorji, Luiz G. Hafemann, Bahareh Pourbabaee, and Mehrsan Javan. Group activity detection from trajectory and video data in soccer. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 898–899, 2020.
[26] Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. Deep high-resolution representation learning for human pose estimation. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5693–5703, 2019.
[27] Masaki Takahashi, Kensuke Ikeya, Masanori Kano, Hidehiko Ookubo, and Tomoyuki Mishina. Robust volleyball tracking system using multi-view cameras. In Proceedings of International Conference on Pattern Recognition, pages 2740–2745, 2016.
[28] Amin Ullah, Khan Muhammad, Javier Del Ser, Sung Wook Baik, and Victor Hugo C. de Albuquerque. Activity recognition using temporal optical flow convolutional features and multilayer lstm. IEEE Transactions on Industrial Electronics, 66(12):9692– 9702, 2018.
[29] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of Advances in Neural Information Processing Systems, pages 5998–6008, 2017.
[30] Roman Voeikov, Nikolay Falaleev, and Ruslan Baikulov. Ttnet: Real-time temporal and spatial video analysis of table tennis. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 884–885, 2020.
[31] Jianchao Wu, Limin Wang, Li Wang, Jie Guo, and Gangshan Wu. Learning actor relation graphs for group activity recognition. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9964–9974, 2019.