| 研究生: |
黃泓茗 Huang, Hung-Ming |
|---|---|
| 論文名稱: |
結構化事件資訊導引排球回合影片描述生成 Structured Event Log-Guided Caption Generation for Volleyball Rally Understanding |
| 指導教授: |
徐禕佑
Hsu, Yi-Yu |
| 學位類別: |
碩士 Master |
| 系所名稱: |
敏求智慧運算學院 - 智慧科技系統碩士學位學程 MS Degree Program on Intelligent Technology Systems |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 79 |
| 中文關鍵詞: | 結構化事件資訊 、視覺語言模型 、排球回合影片理解 、影片描述生成 、幻覺 、LLM-as-a-Judge |
| 外文關鍵詞: | Structured Event Logs, Vision-Language Model, Volleyball Rally Understanding, Video Caption Generation, Hallucination, LLM-as-a-Judge |
| 相關次數: | 點閱:66 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究針對排球比賽影片之回合層級描述生成任務,提出一套結構化事件資訊導引之視覺語言模型(Vision-Language Model, VLM)描述生成框架。排球比賽具有快速攻防轉換、多人互動、時序依賴與規則判斷等特性,通用VLM雖能生成流暢文字,但在排球場景中容易出現動作錯誤、時序錯亂、得分方錯誤與得分原因錯誤等事件層級幻覺。
為改善上述問題,本研究整合既有排球領域電腦視覺模組:裁判手勢辨識模組提供回合切割、發球方與得分方資訊;群體動作偵測模組提供帶時間戳記之隊伍動作序列。此類資訊被整理為結構化事件資訊,作為VLM生成排球回合描述時之輔助上下文。本研究進一步提出WinReasonLogicalReasoning 模組,依動作序列中是否包含攻擊、發球方、得分方與最後攻擊方等資訊,以規則式推論得分原因;該模組屬部分規則式推論,可穩定推論serve_ace、serve_error 與 attack_point,並將難以區分之attack_fault 與 block_point 交由 VLM 依影片判斷。
本研究以2025台灣企排聯賽於成大場地錄製之比賽影片整理136段回合片段進行可行性驗證,透過GPT-4o為評審之G-EVAL式評估、客觀指標(發球方準確率、得分方準確率、得分原因準確率)與排球領域專家評估檢驗生成品質。實驗結果顯示:加入結構化事件資訊後,發球方準確率由0.735提升至0.978、得分方準確率由0.662提升至0.919;再加入推論得分原因後,得分原因準確率由0.588提升至0.838、發球方與得分方準確率提升至0.993,G-EVAL之規則與得分正確性由3.23提升至4.14,逐回合配對檢定均達統計顯著。單一排球領域專家於20段抽樣之成對偏好亦支持逐層加入線索(排除平手後偏好率為0.92與0.94),且以真值事件資訊之消融分析顯示前端誤差傳遞之代價明顯小於已實現之增益(以規則與得分正確性計,實際系統約實現75.9%之可得增益)。實驗亦於多個地端小型VLM上驗證結構化事件資訊之導引效果。上述結果顯示,在本研究之可行性驗證設定下,結合領域CV模組之結構化事件資訊與規則式得分原因推論,可有效降低VLM於排球回合描述生成中之事件層級幻覺。
This thesis studies rally-level caption generation for volleyball match videos. General-purpose Vision-Language Models (VLMs) produce fluent captions but frequently hallucinate at the event level in volleyball scenes, generating incorrect actions, disordered event sequences, wrong scoring sides, and wrong win reasons. This study proposes a structured event log-guided captioning framework that integrates the outputs of existing domain-specific computer vision modules, namely a referee gesture recognition module and a group activity detection module, into structured event logs that serve as grounding context for VLM caption generation. A Win Reason Logical Reasoning module is further proposed to infer the win reason from the event logs through partial deterministic rules. Experiments on 136 rally clips curated from 2025 Taiwan Top Volleyball League matches show that structured event logs substantially improve serving side and scoring side correctness, and that the inferred win reason cue further improves Win Reason Accuracy from 0.588 to 0.838 and the rule and scoring correctness score from 3.23 to 4.14, with per-clip paired tests confirming significance and expert pairwise preferences (0.92 and 0.94) independently supporting each added layer of guidance. The results demonstrate that combining domain-specific structured event information with rule-based reasoning effectively reduces event-level hallucinations in VLM-generated volleyball rally captions.
[1] Joya Chen, Ziyun Zeng, Yiqi Lin, Wei Li, Zejun Ma, and Mike Zheng Shou. LiveCC:Learning video LLM with streaming speech transcription at scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025.
[2] Christoph Feichtenhofer. X3D: Expanding architectures for efficient video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 203–213, 2020.
[3] Gemini Team, Rohan Anil, Sebastian Borgeaud, et al. Gemini: A family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
[4] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022.
[5] Mostafa S. Ibrahim, Srikanth Muralidharan, Zhiwei Deng, Arash Vahdat, and Greg Mori. A hierarchical deep temporal model for group activity recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1971–1980, 2016.
[6] Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, 2023.
[7] Sunghoon Jung, Donghwan Kim, Jihun Park, Hogeon Baek, Jaewoong Lee, Sangwon Yoon, and Sungkwan Cho. Integrated AI system for real-time sports broadcasting: Player behavior, game event recognition, and generative AI commentary in basketball games. Applied Sciences, 15(3):1543, 2025.
[8] Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. Dense captioning events in videos. In Proceedings of the IEEE International Conference on Computer Vision, pages 706–715, 2017.
[9] YanweiLi, ChengyaoWang,andJiayaJia. LLaMA-VID:Animageisworth2tokensin large language models. In European Conference on Computer Vision. Springer Nature Switzerland, 2024.
[10] Chin-Yew Lin. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, 2004.
[11] Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-Eval: NLG evaluation using GPT-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, 2023.
[12] OpenAI. GPT-4o system card. arXiv preprint arXiv:2410.21276, 2024.
[13] Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. BLEU: A method for automatic evaluation of machinetranslation. In Proceedings of the 40thAnnualMeeting of the Association for Computational Linguistics, pages 311–318, 2002.
[14] Mengshi Qi, Yunhong Wang, Annan Li, and Jiebo Luo. Sports video captioning via attentive motion representation and group relationship modeling. IEEE Transactions on Circuits and Systems for Video Technology, 30(8):2617–2633, 2019.
[15] Qwen Team. Qwen3-VL technical report. arXiv preprint arXiv:2511.21631, 2025.
[16] Jiayuan Rao, Haoning Wu, Chang Liu, Yanfeng Wang, and Weidi Xie. MatchTime: Towards automatic soccer game commentary generation. arXiv preprint arXiv:2406.18530, 2024.
[17] AnnaRohrbach, Lisa AnneHendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. Object hallucination in image captioning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4035–4045, 2018.
[18] Jinglin Xu, Guohao Rao, Zeyu Che, YangLiu, Zishuo Yi, Sicheng Zheng, YaoLu, et al. FineSports: A multi-person hierarchical sports video dataset for fine-grained action understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21773–21782, 2024.
[19] Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, and Chunyuan Li. Video instruction tuning with synthetic data. arXiv preprint arXiv:2410.02713, 2024. LLaVA-Video.
[20] YianZhao, WenyuLv,ShangliangXu, JinmanWei, GuanzhongWang, QingqingDang, Yi Liu, and Jie Chen. DETRs beat YOLOs on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16965–16974, 2024.
[21] Zhonghan Zhao, Wenhao Chai, Shengyu Hao, Wenhao Hu, Guanhong Wang, Shidong Cao, Mingli Song, Jenq-Neng Hwang, and Gaoang Wang. A survey of deep learning in sports applications: Perception, comprehension, and decision. IEEE Transactions on Visualization and Computer Graphics, 2025.
[22] LianminZheng,Wei-LinChiang, YingSheng, SiyuanZhuang, ZhanghaoWu,Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems, volume 36, 2023.
[23] Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, et al. InternVL3: Exploring advanced training and test-time recipes for open source multimodal models. arXiv preprint arXiv:2504.10479, 2025.
[24] 施穆然. 排球比賽中的端到端時空即時動作檢測與群體活動識別. Master’sthesis, 國立成功大學人工智慧機器人碩士學位學程,2024. https://hdl.handle.net/11296/q76kgp.
[25] 葉詩棋. 基於深度學習之裁判手勢辨識應用於排球賽事偵測於比賽切片及自動計分. Master’s thesis, 國立成功大學智慧科技系統碩士學位學程, 2025.https://hdl.handle.net/11296/5ba6us.