簡易檢索 / 詳目顯示

研究生: 蕭合亭
Hsiao, Ho-Ting
論文名稱: 基於分層分解圖卷積網路之棒球打擊骨架動作切分及動作分析
Skeleton-Based Motion Segmentation Using Hierarchically Decomposed Graph Convolutional Network and Motion Analysis for Baseball Hitter
指導教授: 連震杰
Lien, Jenn-Jier
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 72
中文關鍵詞: 棒球打擊動作分析 、骨架動作辨識 、關鍵幀偵測 、圖卷積網路 、動作階段分割 、HD-GCN
外文關鍵詞: Baseball Hitter Swing Motion Analysis, Skeleton-Based Action Recognition, Keyframe Detection, Graph Convolution, Temporal Phase Segmentation, HD-GCN
相關次數: 點閱:85  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 精確的棒球打擊動作分析對於提升運動員表現與預防運動傷害至關重要。然而,傳統光學動作捕捉系統設備昂貴、架設繁瑣且需穿戴標記,受限於受控的室內實驗室環境,缺乏實際牛棚訓練之生態效度;而現有基於深度學習的動作辨識模型多僅能輸出單一全局標籤,無法對連續未裁切的打擊過程進行影格層級的細粒度動作階段分割。
    為解決上述挑戰,本研究開發了一套全自動、無標記之三維棒球打擊骨架動作切分與生物力學分析系統。在系統工程方面,本系統整合張氏標定法與八點演算法進行雙相機空間幾何校正,並導入先進先出佇列(FIFO Queue)與事件觸發機制,自動擷取 90 幀之完整連續打擊過程。結合 YOLOv11、ViTPose 與凸包演算法,系統能精準重建人體與球棒之 3D 空間姿態,並量化算出重心速度、擊球仰角及動力鏈傳遞順序(骨盆、軀幹、手臂與球棒)等運動科學診斷指標。
    在演算法創新方面,本研究針對分層分解圖卷積網路(HD-GCN)進行深度優化:(1) 導入滑動視窗機制(Sliding Window),將影片級別分類模型升級為逐幀階段分割框架;(2) 採用焦點損失函數(Focal Loss)取代傳統交叉熵,有效克服擊球點(Impact)極短瞬間所導致的極端類別不平衡問題;(3) 提出邊界填充(Boundary Padding)與時序偏差補償(Bias Correction)策略,在序列末端加入 32 幀 Padding 解決邊界特徵截斷問題,使 Impact 階段的召回率(Recall)顯著提升,並透過物理位移校正消除系統固有滯後延遲。

    Precise baseball swing analysis is crucial for enhancing athletic performance and preventing sports injuries. However, traditional optical motion capture systems require expensive hardware, complex setups, and physical markers, restricting their use to controlled indoor laboratory environments and limiting their ecological validity in bullpen settings. Furthermore, existing deep learning-based action recognition models primarily output a single global class label, failing to achieve fine-grained, frame-level temporal phase segmentation on continuous, untrimmed hitting video sequences.
    To address these challenges, this thesis presents a fully automated, markerless 3D baseball swing segmentation and biomechanical analysis system. In terms of system engineering, the proposed framework integrates Zhang's Calibration, the 8-Point Algorithm, and SVD for robust dual-camera spatial calibration, and incorporates a First-In-First-Out (FIFO) queue with an event-triggered mechanism to autonomously capture a complete 90-frame swing sequence. By combining YOLOv11, ViTPose, and Convex Hull algorithms, the system accurately reconstructs 3D spatial postures of the hitter and the bat, quantifying actionable sports-science metrics such as center-of-mass velocity, attack angle, and kinematic sequence order.
    Regarding algorithmic innovations, the baseline Hierarchically Decomposed Graph Convolutional Network (HD-GCN) is substantially optimized: (1) A Sliding Window Mechanism is incorporated to transform HD-GCN from a video-level classifier into a frame-wise phase segmentation framework; (2) Focal Loss replaces standard Cross-Entropy Loss to dynamically down-weight background samples, mitigating severe class imbalance caused by the extremely brief Impact moment; (3) Boundary Padding and Temporal Bias Correction strategies are introduced.

    摘要 III Abstract IV 誌謝 V List of Tables IX List of Figures X Chapter 1 Introduction 1 1.1 Motivation and Objectives 1 1.2 Global Framework 2 1.3 Related Work 6 1.4 Contributions 11 Chapter 2 System Setup and Specifications 14 2.1 System Setup 14 2.2 Hardware Specifications 15 2.3 User Interface for Display and Labeling 17 Chapter 3 4-Phase Segmentation and Motion Analysis for Baseball Hitter Swing 22 3.1 Definition of Baseball Hitter Swing Motions 22 3.2 Camera Calibration System Architecture 25 3.2.1 Intrinsic and Extrinsic Parameters Calibration 25 3.2.2 Event-Triggered Recording for Video Clips 28 3.2.3 4-Phase Segmentation and Motion Analysis in 2D Mode 29 3.2.4 4-Phase Segmentation and Motion Analysis in 3D Mode 31 3.3 Functions of Hitter’s Swing Analysis 34 3.3.1 Functions of Hitter’s Swing Analysis in 2D Mode 34 3.3.2 Functions of Hitter’s Swing Analysis in 3D Mode 36 Chapter 4 4-Phase Segmentation Based on Keyframe Detection using HD-GCN 38 4.1 Framework of HD-GCN Classification 38 4.2 Framework of HD-GCN Block 42 4.2.1 Graph Structure and Adjacency Matrix Generation of the HD-GCN Spatial Convolution Module 43 4.2.2 HD-Graph Spatial Convolution Block 46 4.2.3 AHA Attention Aggregation Block 47 4.2.4 Temporal Convolution and Down-Sampling Module 48 4.3 Loss Function Design 49 Chapter 5 Experimental Results 51 5.1 Data Collection and Metrics 51 5.1.1 Data Collection 51 5.1.2 Metrics 53 5.2 Experimental Results 54 5.2.1 Performance Evaluation of the Traditional 2D Geometric Coordinate Method 54 5.2.2 Performance of the HD-GCN Model 55 5.2.3 Core Advantages and Theoretical Analysis 55 5.3 Results Analysis 56 Chapter 6 Conclusion and Future Work 58 6.1 Conclusion 58 6.2 Future Work 58 References 60

    [1] B.M. Booth, K. Mundnich, and S.S. Narayanan, “A Novel Method for Human Bias Correction of Continuous-Time Annotations,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 3091-3095, 2018.
    [2] K. Cheng, Y. Zhang, X. He, W. Chen, J. Cheng, and H. Lu, “Skeleton-Based Action Recognition with Shift Graph Convolutional Network,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 183-192, 2020
    [3] R.I. Hartley, “In Defense of the Eight-Point Algorithm,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 19, No. 6, pp. 580-593, 1997.
    [4] R.I. Hartley, and P. Sturm, “Triangulation,” in Computer Vision and Image Understanding, Vol. 68, No. 2, pp.146-157, 1997.
    [5] R. Khanam, and M. Hussain, “YOLOv11: An Overview of the Key Architectural Enhancements,” in arXiv preprint arXiv: 2410.17725, 2024.
    [6] T.Y. Lo, P.Y. Chuang, and J.H. Chang, “Motion Analysis of Baseball Batting from Different Hitting-Point Heights in Different Collegiate Divisions,” in Chinese Journal of Sports Biomechanics, Vol. 16, Issue. 2, 2019.
    [7] C. Lea, M. D. Flynn, R. Vidal, A. Reiter, and G.D. Hager, “Temporal Convolutional Networks for Action Segmentation and Detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 156-165, 2017
    [8] T.Y. Lin, R. Goyal, R. Girshick, K. He, and P. Dollar, “Focal Loss for Dense Object Detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2980-2988, 2017.
    [9] J. Lee, M. Lee, D. Lee, and S. Lee, “Hierarchically Decomposed Graph Convolutional Networks for Skeleton-Based Action Recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10445-10453, 2023.
    [10] B. Li, X. Li, Z. Zhang, and F. Wu, “Spatio-Temporal Graph Routing for Skeleton-Based Action Recognition,” in Proceedings of the AAAI conference on Artificial Intelligence. Vol. 33, No. 01, 2019.
    [11] W. McNally, K. Vats, T. Pinto, C. Dulhanty, J. McPhee, and A. Wong, “GolfDB: A Video Database for Golf Swing Sequencing,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2019
    [12] Z. Shou, D. Wang, and S.F. Chang, “Temporal Action Localization in Untrimmed Videos via Multi-stage CNNs,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1049-1058, 2016.
    [13] L. Shi, Y. Zhang, J. Cheng, and H. Lu, “Skeleton-Based Action Recognition with Directed Graph Neural Networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7912-7921, 2019.
    [14] C.M. Welch, S.A. Banks, F.F. Cook, and P. Draovitch, “Hitting a Baseball: A Biomechanical Description,” in Journal of Orthopaedic and Sports Physical Therapy, Vol 22., No. 5, pp.193-201, 1995.
    [15] A. Wang, H. Chen, L. Liu, K. Chen, Z. Lin, J. Han, and G. Ding, “YOLOv10: Real-Time End-to-End Object Detection,” Advances in Neural Information Processing Systems, Vol. 37, pp. 107984-108011, 2024.
    [16] N. Wiedemann, C. Dietrich, and C.T. Silva, “A Tracking System for Baseball Game Reconstruction,” in arXiv preprint arXiv:2003.03856, 2020.
    [17] L. Wang, Y. Xiong, Y. Qiao, D. Lin, X. Tang, and L.V. Gool, “Temporal Segment Networks for Action Recognition in Videos,” in Proceedings of IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol.41, No.11, pp. 2740-2755, 2019.
    [18] Y. Xu, J. Zhang, Q. Zhang, and D. Tao, “ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation,” Advances in Neural Information Processing Systems, Vol. 35, pp. 38571-38584, 2022.
    [19] S. Yan, Y. Xiong, and D. Lin, “Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition,” in Proceedings of the AAAI conference on artificial intelligence, Vol. 32, No. 01, 2018.
    [20] Z. Zhang, “A Flexible New Technique for Camera Calibration,” in Proceedings of IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 22, No. 11, pp. 1330-1334, 2000.
    [21] Y. Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, Y. Liu, and J. Chen, “DETRs Beat YOLOs on Real-time Object Detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16965-16974, 2024.

    下載圖示
    校外:立即公開
    QR CODE