簡易檢索 / 詳目顯示

研究生: 黃泓茗
Huang, Hung-Ming
論文名稱: 結構化事件資訊導引排球回合影片描述生成
Structured Event Log-Guided Caption Generation for Volleyball Rally Understanding
指導教授: 徐禕佑
Hsu, Yi-Yu
學位類別: 碩士
Master
系所名稱: 敏求智慧運算學院 - 智慧科技系統碩士學位學程
MS Degree Program on Intelligent Technology Systems
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 79
中文關鍵詞: 結構化事件資訊視覺語言模型排球回合影片理解影片描述生成幻覺LLM-as-a-Judge
外文關鍵詞: Structured Event Logs, Vision-Language Model, Volleyball Rally Understanding, Video Caption Generation, Hallucination, LLM-as-a-Judge
相關次數: 點閱:66下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究針對排球比賽影片之回合層級描述生成任務,提出一套結構化事件資訊導引之視覺語言模型(Vision-Language Model, VLM)描述生成框架。排球比賽具有快速攻防轉換、多人互動、時序依賴與規則判斷等特性,通用VLM雖能生成流暢文字,但在排球場景中容易出現動作錯誤、時序錯亂、得分方錯誤與得分原因錯誤等事件層級幻覺。
    為改善上述問題,本研究整合既有排球領域電腦視覺模組:裁判手勢辨識模組提供回合切割、發球方與得分方資訊;群體動作偵測模組提供帶時間戳記之隊伍動作序列。此類資訊被整理為結構化事件資訊,作為VLM生成排球回合描述時之輔助上下文。本研究進一步提出WinReasonLogicalReasoning 模組,依動作序列中是否包含攻擊、發球方、得分方與最後攻擊方等資訊,以規則式推論得分原因;該模組屬部分規則式推論,可穩定推論serve_ace、serve_error 與 attack_point,並將難以區分之attack_fault 與 block_point 交由 VLM 依影片判斷。
    本研究以2025台灣企排聯賽於成大場地錄製之比賽影片整理136段回合片段進行可行性驗證,透過GPT-4o為評審之G-EVAL式評估、客觀指標(發球方準確率、得分方準確率、得分原因準確率)與排球領域專家評估檢驗生成品質。實驗結果顯示:加入結構化事件資訊後,發球方準確率由0.735提升至0.978、得分方準確率由0.662提升至0.919;再加入推論得分原因後,得分原因準確率由0.588提升至0.838、發球方與得分方準確率提升至0.993,G-EVAL之規則與得分正確性由3.23提升至4.14,逐回合配對檢定均達統計顯著。單一排球領域專家於20段抽樣之成對偏好亦支持逐層加入線索(排除平手後偏好率為0.92與0.94),且以真值事件資訊之消融分析顯示前端誤差傳遞之代價明顯小於已實現之增益(以規則與得分正確性計,實際系統約實現75.9%之可得增益)。實驗亦於多個地端小型VLM上驗證結構化事件資訊之導引效果。上述結果顯示,在本研究之可行性驗證設定下,結合領域CV模組之結構化事件資訊與規則式得分原因推論,可有效降低VLM於排球回合描述生成中之事件層級幻覺。

    This thesis studies rally-level caption generation for volleyball match videos. General-purpose Vision-Language Models (VLMs) produce fluent captions but frequently hallucinate at the event level in volleyball scenes, generating incorrect actions, disordered event sequences, wrong scoring sides, and wrong win reasons. This study proposes a structured event log-guided captioning framework that integrates the outputs of existing domain-specific computer vision modules, namely a referee gesture recognition module and a group activity detection module, into structured event logs that serve as grounding context for VLM caption generation. A Win Reason Logical Reasoning module is further proposed to infer the win reason from the event logs through partial deterministic rules. Experiments on 136 rally clips curated from 2025 Taiwan Top Volleyball League matches show that structured event logs substantially improve serving side and scoring side correctness, and that the inferred win reason cue further improves Win Reason Accuracy from 0.588 to 0.838 and the rule and scoring correctness score from 3.23 to 4.14, with per-clip paired tests confirming significance and expert pairwise preferences (0.92 and 0.94) independently supporting each added layer of guidance. The results demonstrate that combining domain-specific structured event information with rule-based reasoning effectively reduces event-level hallucinations in VLM-generated volleyball rally captions.

    摘要 i Abstract ii 英文延伸摘要 iv TableofContents viii ListofTables xi ListofFigures xii Nomenclature xiii Chapter1.緒論 1 1.1.研究背景 1 1.2.研究動機 1 1.3.研究問題 2 1.4.研究目的 3 1.5.研究貢獻 3 1.6.論文架構 4 Chapter2.文獻探討 5 2.1.運動影片理解 5 2.2.排球事件偵測與動作辨識 5 2.3. VideoCaptioning與DenseVideoCaptioning 6 2.4.視覺語言模型於影片理解之應用 6 2.5. VLMHallucination與Grounding問題 7 2.6. LLM-as-a-Judge與生成文字評估 7 2.7.小結 7 Chapter3.研究資料與前端事件資訊抽取 9 3.1.系統整體流程概述 9 3.2.資料來源與回合片段 10 3.3.人工標註內容 11 3.4. WinReason類別定義與分布 12 3.5.裁判手勢辨識模組 13 3.6.群體動作偵測模組 14 3.7.結構化事件資訊建構 15 Chapter4.結構化事件資訊導引之描述生成方法 17 4.1.方法總覽 17 4.2. Video-OnlyCaptionGenerationBaseline 18 4.3. Event-GuidedCaptionGeneration 18 4.4. WinReasonLogicalReasoning 19 4.4.1.設計動機 19 4.4.2.推論流程 19 4.4.3.模組性質與嚴謹性說明 21 4.4.4.推論結果之使用方式 21 4.5. EnhancedEvent-GuidedCaptionGeneration 21 4.6. PromptDesign 22 4.7. Qwen3-VL-8BLoRAFine-Tuning 23 4.8.應用展示:Broadcast-StyleCommentaryGeneration 23 Chapter5.實驗設計與結果分析 26 5.1.實驗資料與設定 26 5.2.評估指標 27 5.2.1. G-EVAL/LLM-as-a-Judge 27 5.2.2.客觀指標 28 5.2.3.專家評估 28 5.3.前端事件資訊抽取之準確率 29 5.4.主實驗:結構化事件資訊與推論得分原因之效果 30 5.4.1. G-EVAL結果 30 5.4.2.逐回合配對統計檢定 31 5.4.3.客觀指標結果 32 5.4.4.得分原因之混淆矩陣分析 33 5.5. GT事件資訊消融:誤差傳遞與效能上限 33 5.6. WinReasonLogicalReasoning推論準確率 35 5.7.地端小型VLM之比較 36 5.8. Qwen3-VL-8BLoRA微調結果 36 5.9.專家評估結果 37 5.9.1.評估設定 37 5.9.2.專家評分結果 38 5.9.3.成對偏好結果 39 5.9.4.小結與限制 39 5.10.應用比較:播報風格解說與LiveCC 40 5.11. CaseStudy 41 5.12.討論與限制 42 5.12.1.主要發現 42 5.12.2.研究限制 43 Chapter6.結論與未來工作 45 6.1.研究結論 45 6.2.研究貢獻回顧 46 6.3.研究限制 46 6.4.未來工作 47 References 49 AppendixA.完整CaptionGenerationPrompt 52 A.1. Video-Only設定 52 A.2. Event-Guided設定 52 A.3. Enhanced設定之附加指示 54 AppendixB. G-EVALJudgePrompt 55 AppendixC. ObjectiveMetricExtractionPrompt 59 AppendixD.生成結果案例 61 D.1. Video-Only設定之生成結果 61 D.2. Event-Guided設定之生成結果 61 D.3. Enhanced設定之生成結果 62 AppendixE.專家評估表單與評分結果 63 E.1.評估表單 63 E.2.專家評分之描述統計 63 E.3.專家與LLM評審之平均分數對照 64

    [1] Joya Chen, Ziyun Zeng, Yiqi Lin, Wei Li, Zejun Ma, and Mike Zheng Shou. LiveCC:Learning video LLM with streaming speech transcription at scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025.
    [2] Christoph Feichtenhofer. X3D: Expanding architectures for efficient video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 203–213, 2020.
    [3] Gemini Team, Rohan Anil, Sebastian Borgeaud, et al. Gemini: A family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
    [4] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022.
    [5] Mostafa S. Ibrahim, Srikanth Muralidharan, Zhiwei Deng, Arash Vahdat, and Greg Mori. A hierarchical deep temporal model for group activity recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1971–1980, 2016.
    [6] Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, 2023.
    [7] Sunghoon Jung, Donghwan Kim, Jihun Park, Hogeon Baek, Jaewoong Lee, Sangwon Yoon, and Sungkwan Cho. Integrated AI system for real-time sports broadcasting: Player behavior, game event recognition, and generative AI commentary in basketball games. Applied Sciences, 15(3):1543, 2025.
    [8] Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. Dense captioning events in videos. In Proceedings of the IEEE International Conference on Computer Vision, pages 706–715, 2017.
    [9] YanweiLi, ChengyaoWang,andJiayaJia. LLaMA-VID:Animageisworth2tokensin large language models. In European Conference on Computer Vision. Springer Nature Switzerland, 2024.
    [10] Chin-Yew Lin. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, 2004.
    [11] Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-Eval: NLG evaluation using GPT-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, 2023.
    [12] OpenAI. GPT-4o system card. arXiv preprint arXiv:2410.21276, 2024.
    [13] Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. BLEU: A method for automatic evaluation of machinetranslation. In Proceedings of the 40thAnnualMeeting of the Association for Computational Linguistics, pages 311–318, 2002.
    [14] Mengshi Qi, Yunhong Wang, Annan Li, and Jiebo Luo. Sports video captioning via attentive motion representation and group relationship modeling. IEEE Transactions on Circuits and Systems for Video Technology, 30(8):2617–2633, 2019.
    [15] Qwen Team. Qwen3-VL technical report. arXiv preprint arXiv:2511.21631, 2025.
    [16] Jiayuan Rao, Haoning Wu, Chang Liu, Yanfeng Wang, and Weidi Xie. MatchTime: Towards automatic soccer game commentary generation. arXiv preprint arXiv:2406.18530, 2024.
    [17] AnnaRohrbach, Lisa AnneHendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. Object hallucination in image captioning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4035–4045, 2018.
    [18] Jinglin Xu, Guohao Rao, Zeyu Che, YangLiu, Zishuo Yi, Sicheng Zheng, YaoLu, et al. FineSports: A multi-person hierarchical sports video dataset for fine-grained action understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21773–21782, 2024.
    [19] Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, and Chunyuan Li. Video instruction tuning with synthetic data. arXiv preprint arXiv:2410.02713, 2024. LLaVA-Video.
    [20] YianZhao, WenyuLv,ShangliangXu, JinmanWei, GuanzhongWang, QingqingDang, Yi Liu, and Jie Chen. DETRs beat YOLOs on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16965–16974, 2024.
    [21] Zhonghan Zhao, Wenhao Chai, Shengyu Hao, Wenhao Hu, Guanhong Wang, Shidong Cao, Mingli Song, Jenq-Neng Hwang, and Gaoang Wang. A survey of deep learning in sports applications: Perception, comprehension, and decision. IEEE Transactions on Visualization and Computer Graphics, 2025.
    [22] LianminZheng,Wei-LinChiang, YingSheng, SiyuanZhuang, ZhanghaoWu,Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems, volume 36, 2023.
    [23] Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, et al. InternVL3: Exploring advanced training and test-time recipes for open source multimodal models. arXiv preprint arXiv:2504.10479, 2025.
    [24] 施穆然. 排球比賽中的端到端時空即時動作檢測與群體活動識別. Master’sthesis, 國立成功大學人工智慧機器人碩士學位學程,2024. https://hdl.handle.net/11296/q76kgp.
    [25] 葉詩棋. 基於深度學習之裁判手勢辨識應用於排球賽事偵測於比賽切片及自動計分. Master’s thesis, 國立成功大學智慧科技系統碩士學位學程, 2025.https://hdl.handle.net/11296/5ba6us.

    下載圖示
    校外:立即公開
    QR CODE