| 研究生: |
張光德 Chang, Kuang-Te |
|---|---|
| 論文名稱: |
應用於人體動作預測之量子殘差學習 Quantum Residual Learning for Human Motion Prediction |
| 指導教授: |
林家祥
Lin, Chia-Hsiang |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電腦與通信工程研究所 Institute of Computer & Communication Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 42 |
| 中文關鍵詞: | 人體動作預測 、量子機器學習 、殘差學習 、深度學習 |
| 外文關鍵詞: | Human Motion Prediction, Quantum Machine Learning, Residual Learning, Deep Learning |
| 相關次數: | 點閱:82 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
人體動作預測旨在根據歷史骨架序列推估未來的人體姿態,廣泛應用於機器人、人機互動、動畫製作及智慧運動等領域。然而,人體運動包含複雜的時空依賴關係、生物力學限制與不確定性,使得模型在較遠的預測時間範圍內容易產生較大的姿態偏差。因此,如何在保留物理先驗所提供之基礎軌跡的同時,進一步提升中長期預測準確度,仍是一項具有挑戰性的研究課題。
為改善上述問題,本文提出一套量子殘差框架,將凍結參數的PhysMoP骨幹網路與可學習的量子殘差模組進行整合。PhysMoP根據歷史動作序列產生具物理資訊導向的未來基礎預測;接著,將歷史觀測序列與PhysMoP所預測的未來軌跡串接,以形成完整的時序脈絡。量子殘差模組則根據此時序資訊,估計用於修正未來軌跡的殘差特徵,並透過可學習的縮放參數控制殘差修正對最終預測結果的影響。
所提出的量子殘差模組使用參數化量子電路作為另一種特徵轉換機制。時序特徵經線性轉換與區塊分割後,被編碼至量子態,並透過參數化旋轉閘、糾纏操作及量測轉換為殘差特徵。本文並非以量子模組取代原有的物理資訊導向預測器,而是將量子特徵轉換整合至殘差學習架構中,為基礎預測提供互補的動作修正資訊。
為驗證所提出方法,本文於Human3.6M、AMASS-BMLrub與3DPW等人體動作預測基準資料集上進行實驗。相較於凍結的PhysMoP骨幹網路,所提出框架在多數評估時間點皆獲得較低的平均每關節位置誤差,且改善主要出現在中期與長期預測範圍,短期預測則維持與PhysMoP相近的表現。AMASS至3DPW的跨資料集實驗亦顯示,量子殘差分支能在未見過的動作分布下提供有效的修正資訊。此外,與傳統MLP-128 殘差模型相比,所提出的量子殘差模組僅使用13.3K個可訓練參數,即可獲得相近的預測準確率;相較於MLP-128的26.2K個參數,參數量減少49.3%。
Human motion prediction is an important task in computer vision, robotics, animation, and human–computer interaction. However, long-term prediction remains challenging because human motion involves complex spatio-temporal dependencies, biomechanical constraints, and increasing uncertainty at distant forecasting horizons. Although PhysMoP provides use ful physical priors, its physics-informed formulationandlearnedcomponentsmaynotcapture all fine-grained motion variations.
To address this issue, we propose a Quantum Residual Framework that combines a frozen PhysMoP backbone with a learnable Quantum Residual Module. PhysMoP first generates a physics-informed future trajectory. The observed history and the predicted future are then concatenated to form a complete temporal context, from which the quantum residual module estimates corrections for refining the future motion.
The proposed module uses parameterized quantum circuits as an alternative feature transformation mechanism. Temporal features are partitioned into local sub-blocks, encoded into quantum states, and processed using parameterized rotation and entanglement operations. Rather than replacing the physics-informed predictor, the quantum module is used to learn complementary residual corrections.
Experiments on Human3.6M, AMASS-BMLrub, and 3DPW demonstrate that the proposed framework reduces MPJPE relative to the frozen PhysMoP backbone across most evaluated prediction horizons. The improvements are most evident in mid-term and long term prediction, while short-term performance remains comparable to the original backbone. Cross-dataset evaluation from AMASS to 3DPW also shows improvements at multiple prediction horizons without additional fine-tuning. Compared with the classical MLP-128 residual refiner, the proposed quantum residual module achieves comparable prediction accuracy using 13.3K trainable parameters, corresponding to a 49.3% reduction relative to the 26.2K parameters of the MLP-128 refiner.
[1] J. Martinez, M. J. Black, and J. Romero, “On human motion prediction using recurrent neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2891–2900.
[2] W. Mao, M. Liu, M. Salzmann, and H. Li, “Learning trajectory dependencies for human motion prediction,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9489–9497.
[3] W. Mao, M. Liu, and M. Salzmann, “History repeats itself: Human motion prediction via motion attention,” in European Conference on Computer Vision. Springer, 2020, pp. 474–489.
[4] W. Guo, Y. Du, X. Shen, V. Lepetit, X. Alameda-Pineda, and F. Moreno-Noguer, “Back to mlp: A simple baseline for human motion prediction,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2023, pp. 4809–4819.
[5] T. Ma, Y. Nie, C. Long, Q. Zhang, and G. Li, “Progressively generating better initial guesses towards next stages for high-quality human motion prediction,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 6437–6446.
[6] D. Wei, H. Sun, B. Li, J. Lu, W. Li, X. Sun, and S. Hu, “Human joint kinematics diffusion-refinement for stochastic motion prediction,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 5, 2023, pp. 6110–6118.
[7] J. Sun and G. Chowdhary, “CoMusion: Towards consistent stochastic human motion prediction via motion diffusion,” in European Conference on Computer Vision. Springer, 2024, pp. 18–36.
[8] S. Saadatnejad, A. Rasekh, M. Mofayezi, Y. Medghalchi, S. Rajabzadeh, T. Mordan, and A. Alahi, “A generic diffusion-based approach for 3d human pose prediction in the wild,” in IEEE International Conference on Robotics and Automation, 2023, pp. 8246–8253.
[9] Y. Zhang, J. O. Kephart, and Q. Ji, “Incorporating physics principles for precise human motion prediction,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 6164–6174.
[10] J. Biamonte, P. Wittek, N. Pancotti, P. Rebentrost, N. Wiebe, and S. Lloyd, “Quantum machine learning,” Nature, vol. 549, no. 7671, pp. 195–202, 2017.
[11] P.-W. Tang, C.-H. Lin, J.-K. Huang, and A. R. Huete, “A quantum-empowered spei drought forecasting algorithm using spatially-aware mamba network,” IEEE Transactions on Geoscience and Remote Sensing, 2025.
[12] J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven, “Barren plateaus in quantum neural network training landscapes,” Nature communications, vol. 9, no. 1, p. 4812, 2018.
[13] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
[14] C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu, “Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments,” IEEE transactions on pattern analysis and machine intelligence, vol. 36, no. 7, pp. 1325–1339, 2013.
[15] N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black, “Amass: Archive of motion capture as surface shapes,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 5442–5451.
[16] T. Von Marcard, R. Henschel, M. J. Black, B. Rosenhahn, and G. Pons-Moll,“Recovering accurate 3d human pose in the wild using imus and a moving camera,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 601–617.
[17] G. Moon, H. Choi, and K. M. Lee, “Neuralannot: Neural annotator for 3d human mesh training sets,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 2299–2307.
[18] R. Ji, C. Lu, Z. Huang, and J. Zhong, “Learning behavior aware features across spaces for improved 3d human motion prediction,” Scientific Reports, vol. 15, no. 1, p. 28355, 2025.
[19] Y. Li, L. Zhu, H. Xie, and X. Yi, “Multi-granularity spatiotemporal fusion neural network with denoising reconstruction for human motion prediction,” Applied Soft Computing, p. 113909, 2025.