簡易檢索 / 詳目顯示

研究生: 陳喬雅
Chen, Chiao-Ya
論文名稱: 多視角三維人體姿態估計之根節點引導粗到細關節定位方法
Root-Guided Coarse-to-Fine Joint Localization for Multi-View 3D Human Pose Estimation
指導教授: 蔡家齊
Tsai, Chia-Chi
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 79
中文關鍵詞: 多視角多人三維人體姿態估計 、voxel-based 方法 、one-stage 姿態偵測 、粗到細關節定位
外文關鍵詞: multi-view multi-person 3D human pose estimation, voxel-based methods, one-stage detector, coarse-to-fine joint localization
相關次數: 點閱:85  下載:4 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 多視角多人三維人體姿態估測旨在從同步拍攝的多視角影像中,重建場景內所有人的三維姿態。Voxel-based方法將影像特徵投影至共享的三維空間,但對完整場景進行密集voxel運算會帶來高昂的計算成本。因此,多數方法採用兩階段的由粗至細架構:第一階段先利用較粗解析度的voxel產生人物候選區域,第二階段再於各候選區域內進行細緻的人體姿態估測。雖然此設計降低了完整場景的處理成本,但第二階段仍須針對每一位偵測到的人物重複執行,導致推論複雜度與場景人數相關。
    為解決人數相依的計算問題,本研究提出一種單階段voxel-based三維人體姿態估測方法。首先,將多視角二維關節heatmaps投影至共享的三維特徵體積。接著,透過關節定位網路直接預測三維關節heatmaps,並利用基於人體骨架的匈牙利關節分配演算法,根據根節點資訊進行粗略關節的選取。最後,透過局部特徵擷取與 soft-argmax,取得最終細化後的三維人體關節座標。此設計以共享的由粗至細關節定位流程,取代傳統針對各人物候選區域分別執行的姿態迴歸。
    實驗結果顯示,所提出的方法能有效消除人數相依的計算成本,並在場景人數增加時維持近乎固定的每秒幀數與推論延遲。在 Panoptic 資料集上,本方法僅需 96.24 G MACs,為所有比較方法中計算成本最低者,並達到 18.35 mm 的平均每關節位置誤差。

    Multi-view multi-person 3D human pose estimation recovers the 3D poses of all people from synchronized multi-view images. Voxel-based methods project image features into a shared 3D space, but dense full-scene voxel processing is computationally expensive. Therefore, most methods adopt a two-stage coarse-to-fine architecture: a coarse voxel grid first generates person proposals, followed by fine-grained pose estimation within each proposal. Although this reduces full-scene processing costs, the second stage must be repeated for every detected person, resulting in person-dependent inference complexity. To address person-dependent computation, we proposed a one-stage voxel-based 3D human pose estimation method. Multi-view 2D joint heatmaps are projected into a shared 3D feature volume. A Joint Localization Network then directly predicts joint heatmaps and a Skeleton Hungarian Joint Assignment Algorithm uses root-channel information to associate and select coarse joints. Finally, local feature extraction and soft-argmax are applied to obtain the final refined 3D human joint coordinates. This design replaces proposal-wise pose regression with shared coarse-to-fine joint localization. Experimental results show that the proposed method eliminates person-dependent computation, maintaining nearly constant FPS and latency as the number of people increases. On Panoptic, it achieves 18.35 mm MPJPE with only 96.24 G MACs, the lowest computational cost among the compared methods.

    摘要 iii Abstract v Content vii List of Tables ix List of Figures x Chapter 1 Introduction 1 1.1 Research Background and Motivation 1 1.2 Proposed Method Overview 4 1.3 Experimental Findings and Contributions 6 1.4 Thesis Organization 9 Chapter 2 Related work 10 2.1 Geometry-based reconstruction methods 12 2.1.1 Overview 12 2.1.2 Fast and Robust 13 2.2 Explicit volumetric representation methods 15 2.2.1 Overview 15 2.2.2 VoxelPose 16 2.2.3 Faster VoxelPose 18 2.2.4 3DSA 21 2.2.5 3D Space Token Swinformer 23 2.2.6 Other Subsequent Voxel-based Developments 25 2.3 Implicit representation learning methods 29 2.3.1 Overview 29 2.3.2 MvP 30 2.3.3 MVGFormer 32 Chapter 3 Method 35 3.1 Overview 35 3.2 Multi-View Heatmap Extraction 36 3.3 Coarse-to-Fine Joint Localization 37 3.3.1 Coarse Human Joint Detection 39 3.3.2 Fine Human Joint Refinement 43 3.4 Training Objective 44 Chapter 4 Experiments 46 4.1 Implementation details 46 4.2 Datasets and Evaluation Metrics 48 4.2.1 Panoptic 48 4.2.2 Campus and Shelf 49 4.3 Comparison with Existing Methods on the Panoptic Dataset 50 4.4 Comparison with Existing Method on Shelf and Campus datasets 52 4.5 Ablation Studies 54 4.5.1 Effect of Heatmap Quality and 2D Backbone 54 4.5.2 Effect of the Number of Cameras 56 4.5.3 Feature Volume Resolution Study 57 4.5.4 Inference Speed Analysis 60 Chapter 5 Conclusions and Future Work 62 5.1 Conclusions 62 5.2 Future Work 63 References 65

    [1] Zheng, C., Wu, W., Chen, C., Yang, T., Zhu, S., Shen, J., ... & Shah, M. (2023). Deep learning-based human pose estimation: A survey. ACM computing surveys, 56(1), 1-37.
    [2] Bridgeman, L., Volino, M., Guillemaut, J. Y., & Hilton, A. (2019). Multi-person 3d pose estimation and tracking in sports. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops (pp. 0-0).
    [3] Yang, A., Liu, G., Naeem, W., Wu, D., Zhou, Y., & Chen, L. (2023). A monocular 3D human pose estimation approach for virtual character skeleton retargeting. Journal of Ambi-ent Intelligence and Humanized Computing, 14(7), 9563-9574.
    [4] Tu, H., Wang, C., & Zeng, W. (2020, August). Voxelpose: Towards multi-camera 3d human pose estimation in wild environment. In European conference on computer vision (pp. 197-212). Cham: Springer International Publishing.
    [5] Ye, H., Zhu, W., Wang, C., Wu, R., & Wang, Y. (2022, October). Faster voxelpose: Re-al-time 3d human pose estimation by orthographic projection. In European conference on computer vision (pp. 142-159). Cham: Springer Nature Switzerland.
    [6] Zhang, J., Cai, Y., Yan, S., & Feng, J. (2021). Direct multi-view multi-person 3d pose estimation. Advances in neural information processing systems, 34, 13153-13164.
    [7] Liao, Z., Zhu, J., Wang, C., Hu, H., & Waslander, S. L. (2024). Multiple view geometry transformers for 3d human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 708-717).
    [8] Dong, J., Jiang, W., Huang, Q., Bao, H., & Zhou, X. (2019). Fast and robust mul-ti-person 3D pose estimation from multiple views. In Proceedings of the IEEE/CVF Con-ference on Computer Vision and Pattern Recognition (pp. 7792-7801).
    [9] Huang, C., Jiang, S., Li, Y., Zhang, Z., Traish, J., Deng, C., ... & Da Xu, R. Y. (2020, August). End-to-end dynamic matching network for multi-view multi-person 3d pose estima-tion. In European Conference on Computer Vision (pp. 477-493). Cham: Springer Interna-tional Publishing.
    [10] Chen, H., Guo, P., Li, P., Lee, G. H., & Chirikjian, G. (2020, August). Multi-person 3d pose estimation in crowded scenes based on multi-view geometry. In European Conference on Computer Vision (pp. 541-557). Cham: Springer International Publishing.
    [11] Choudhury, R., Kitani, K. M., & Jeni, L. A. (2023). Tempo: Efficient multi-view pose estimation, tracking, and forecasting. In Proceedings of the IEEE/CVF International Con-ference on Computer Vision (pp. 14750-14760).
    [12] Chen, B. H., & Tsai, C. C. (2024, September). 3DSA: Multi-view 3D Human Pose Es-timation With 3D Space Attention Mechanisms. In European Conference on Computer Vi-sion (pp. 323-339). Cham: Springer Nature Switzerland.
    [13] Srivastav, V., Chen, K., & Padoy, N. (2024). Selfpose3d: Self-supervised multi-person multi-view 3d pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2502-2512).
    [14] 3D Space Token Swinformer : Learning Critical Regions with Window-based Token Pruning in 3D Swin Transformer for Multi-View 3D Human Pose Estimation https://hdl.handle.net/11296/v2de2f
    [15] Ershadi-Nasab, S., Noury, E., Kasaei, S., & Sanaei, E. (2018). Multiple human 3d pose estimation from multiview images. Multimedia Tools and Applications, 77(12), 15573-15601.
    [16] Belagiannis, V., Amin, S., Andriluka, M., Schiele, B., Navab, N., & Ilic, S. (2015). 3d pictorial structures revisited: Multiple human pose estimation. IEEE transactions on pattern analysis and machine intelligence, 38(10), 1929-1942.
    [17] Lin, J., & Lee, G. H. (2021). Multi-view multi-person 3d pose estimation with plane sweep stereo. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 11886-11895).
    [18] Xiao, B., Wu, H., & Wei, Y. (2018). Simple baselines for human pose estimation and tracking. In Proceedings of the European conference on computer vision (ECCV) (pp. 466-481).
    [19] Zhang, Y., Wang, C., Wang, X., Liu, W., & Zeng, W. (2022). Voxeltrack: Multi-person 3d human pose estimation and tracking in the wild. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(2), 2613-2626.
    [20] Reddy, N. D., Guigues, L., Pishchulin, L., Eledath, J., & Narasimhan, S. G. (2021). Tessetrack: End-to-end learnable multi-person articulated 3d pose tracking. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 15190-15200).
    [21] Chen, Y., Gu, R., Huang, O., & Jia, G. (2023). VTP: volumetric transformer for mul-ti-view multi-person 3D pose estimation: Yuxing, Chen et al. Applied Intelligence, 53(22), 26568-26579.
    [22] Iskakov, K., Burkov, E., Lempitsky, V., & Malkov, Y. (2019). Learnable triangulation of human pose. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 7718-7727).
    [23] Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., ... & Guo, B. (2021). Swin trans-former: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 10012-10022).
    [24] Sun, X., Xiao, B., Wei, F., Liang, S., & Wei, Y. (2018). Integral human pose regression. In Proceedings of the European conference on computer vision (ECCV) (pp. 529-545).

    下載圖示
    校外:立即公開
    QR CODE