| 研究生: |
陳喬雅 Chen, Chiao-Ya |
|---|---|
| 論文名稱: |
多視角三維人體姿態估計之根節點引導粗到細關節定位方法 Root-Guided Coarse-to-Fine Joint Localization for Multi-View 3D Human Pose Estimation |
| 指導教授: |
蔡家齊
Tsai, Chia-Chi |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 79 |
| 中文關鍵詞: | 多視角多人三維人體姿態估計 、voxel-based 方法 、one-stage 姿態偵測 、粗到細關節定位 |
| 外文關鍵詞: | multi-view multi-person 3D human pose estimation, voxel-based methods, one-stage detector, coarse-to-fine joint localization |
| 相關次數: | 點閱:85 下載:4 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
多視角多人三維人體姿態估測旨在從同步拍攝的多視角影像中,重建場景內所有人的三維姿態。Voxel-based方法將影像特徵投影至共享的三維空間,但對完整場景進行密集voxel運算會帶來高昂的計算成本。因此,多數方法採用兩階段的由粗至細架構:第一階段先利用較粗解析度的voxel產生人物候選區域,第二階段再於各候選區域內進行細緻的人體姿態估測。雖然此設計降低了完整場景的處理成本,但第二階段仍須針對每一位偵測到的人物重複執行,導致推論複雜度與場景人數相關。
為解決人數相依的計算問題,本研究提出一種單階段voxel-based三維人體姿態估測方法。首先,將多視角二維關節heatmaps投影至共享的三維特徵體積。接著,透過關節定位網路直接預測三維關節heatmaps,並利用基於人體骨架的匈牙利關節分配演算法,根據根節點資訊進行粗略關節的選取。最後,透過局部特徵擷取與 soft-argmax,取得最終細化後的三維人體關節座標。此設計以共享的由粗至細關節定位流程,取代傳統針對各人物候選區域分別執行的姿態迴歸。
實驗結果顯示,所提出的方法能有效消除人數相依的計算成本,並在場景人數增加時維持近乎固定的每秒幀數與推論延遲。在 Panoptic 資料集上,本方法僅需 96.24 G MACs,為所有比較方法中計算成本最低者,並達到 18.35 mm 的平均每關節位置誤差。
Multi-view multi-person 3D human pose estimation recovers the 3D poses of all people from synchronized multi-view images. Voxel-based methods project image features into a shared 3D space, but dense full-scene voxel processing is computationally expensive. Therefore, most methods adopt a two-stage coarse-to-fine architecture: a coarse voxel grid first generates person proposals, followed by fine-grained pose estimation within each proposal. Although this reduces full-scene processing costs, the second stage must be repeated for every detected person, resulting in person-dependent inference complexity. To address person-dependent computation, we proposed a one-stage voxel-based 3D human pose estimation method. Multi-view 2D joint heatmaps are projected into a shared 3D feature volume. A Joint Localization Network then directly predicts joint heatmaps and a Skeleton Hungarian Joint Assignment Algorithm uses root-channel information to associate and select coarse joints. Finally, local feature extraction and soft-argmax are applied to obtain the final refined 3D human joint coordinates. This design replaces proposal-wise pose regression with shared coarse-to-fine joint localization. Experimental results show that the proposed method eliminates person-dependent computation, maintaining nearly constant FPS and latency as the number of people increases. On Panoptic, it achieves 18.35 mm MPJPE with only 96.24 G MACs, the lowest computational cost among the compared methods.
[1] Zheng, C., Wu, W., Chen, C., Yang, T., Zhu, S., Shen, J., ... & Shah, M. (2023). Deep learning-based human pose estimation: A survey. ACM computing surveys, 56(1), 1-37.
[2] Bridgeman, L., Volino, M., Guillemaut, J. Y., & Hilton, A. (2019). Multi-person 3d pose estimation and tracking in sports. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops (pp. 0-0).
[3] Yang, A., Liu, G., Naeem, W., Wu, D., Zhou, Y., & Chen, L. (2023). A monocular 3D human pose estimation approach for virtual character skeleton retargeting. Journal of Ambi-ent Intelligence and Humanized Computing, 14(7), 9563-9574.
[4] Tu, H., Wang, C., & Zeng, W. (2020, August). Voxelpose: Towards multi-camera 3d human pose estimation in wild environment. In European conference on computer vision (pp. 197-212). Cham: Springer International Publishing.
[5] Ye, H., Zhu, W., Wang, C., Wu, R., & Wang, Y. (2022, October). Faster voxelpose: Re-al-time 3d human pose estimation by orthographic projection. In European conference on computer vision (pp. 142-159). Cham: Springer Nature Switzerland.
[6] Zhang, J., Cai, Y., Yan, S., & Feng, J. (2021). Direct multi-view multi-person 3d pose estimation. Advances in neural information processing systems, 34, 13153-13164.
[7] Liao, Z., Zhu, J., Wang, C., Hu, H., & Waslander, S. L. (2024). Multiple view geometry transformers for 3d human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 708-717).
[8] Dong, J., Jiang, W., Huang, Q., Bao, H., & Zhou, X. (2019). Fast and robust mul-ti-person 3D pose estimation from multiple views. In Proceedings of the IEEE/CVF Con-ference on Computer Vision and Pattern Recognition (pp. 7792-7801).
[9] Huang, C., Jiang, S., Li, Y., Zhang, Z., Traish, J., Deng, C., ... & Da Xu, R. Y. (2020, August). End-to-end dynamic matching network for multi-view multi-person 3d pose estima-tion. In European Conference on Computer Vision (pp. 477-493). Cham: Springer Interna-tional Publishing.
[10] Chen, H., Guo, P., Li, P., Lee, G. H., & Chirikjian, G. (2020, August). Multi-person 3d pose estimation in crowded scenes based on multi-view geometry. In European Conference on Computer Vision (pp. 541-557). Cham: Springer International Publishing.
[11] Choudhury, R., Kitani, K. M., & Jeni, L. A. (2023). Tempo: Efficient multi-view pose estimation, tracking, and forecasting. In Proceedings of the IEEE/CVF International Con-ference on Computer Vision (pp. 14750-14760).
[12] Chen, B. H., & Tsai, C. C. (2024, September). 3DSA: Multi-view 3D Human Pose Es-timation With 3D Space Attention Mechanisms. In European Conference on Computer Vi-sion (pp. 323-339). Cham: Springer Nature Switzerland.
[13] Srivastav, V., Chen, K., & Padoy, N. (2024). Selfpose3d: Self-supervised multi-person multi-view 3d pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2502-2512).
[14] 3D Space Token Swinformer : Learning Critical Regions with Window-based Token Pruning in 3D Swin Transformer for Multi-View 3D Human Pose Estimation https://hdl.handle.net/11296/v2de2f
[15] Ershadi-Nasab, S., Noury, E., Kasaei, S., & Sanaei, E. (2018). Multiple human 3d pose estimation from multiview images. Multimedia Tools and Applications, 77(12), 15573-15601.
[16] Belagiannis, V., Amin, S., Andriluka, M., Schiele, B., Navab, N., & Ilic, S. (2015). 3d pictorial structures revisited: Multiple human pose estimation. IEEE transactions on pattern analysis and machine intelligence, 38(10), 1929-1942.
[17] Lin, J., & Lee, G. H. (2021). Multi-view multi-person 3d pose estimation with plane sweep stereo. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 11886-11895).
[18] Xiao, B., Wu, H., & Wei, Y. (2018). Simple baselines for human pose estimation and tracking. In Proceedings of the European conference on computer vision (ECCV) (pp. 466-481).
[19] Zhang, Y., Wang, C., Wang, X., Liu, W., & Zeng, W. (2022). Voxeltrack: Multi-person 3d human pose estimation and tracking in the wild. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(2), 2613-2626.
[20] Reddy, N. D., Guigues, L., Pishchulin, L., Eledath, J., & Narasimhan, S. G. (2021). Tessetrack: End-to-end learnable multi-person articulated 3d pose tracking. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 15190-15200).
[21] Chen, Y., Gu, R., Huang, O., & Jia, G. (2023). VTP: volumetric transformer for mul-ti-view multi-person 3D pose estimation: Yuxing, Chen et al. Applied Intelligence, 53(22), 26568-26579.
[22] Iskakov, K., Burkov, E., Lempitsky, V., & Malkov, Y. (2019). Learnable triangulation of human pose. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 7718-7727).
[23] Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., ... & Guo, B. (2021). Swin trans-former: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 10012-10022).
[24] Sun, X., Xiao, B., Wei, F., Liang, S., & Wei, Y. (2018). Integral human pose regression. In Proceedings of the European conference on computer vision (ECCV) (pp. 529-545).