簡易檢索 / 詳目顯示

研究生: 許志煒
Xu, Zhi-Wei
論文名稱: 基於殘差時間標準化與混合時頻監督的光照強健遠程心率估計
Illumination-Robust Remote Heart Rate Estimation with Residual Temporal Standardization and Hybrid Temporal-Frequency Supervision
指導教授: 吳馬丁
Nordling, Torbjörn E.M.
學位類別: 碩士
Master
系所名稱: 工學院 - 機械工程學系
Department of Mechanical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 82
中文關鍵詞: 遠端光體積變化描記法心率估測光照穩健性時間域與頻率域混合損失函數殘差式時間標準化模組
外文關鍵詞: remote photoplethysmography, heart rate estimation, illumination robustness, temporal-frequency hybrid loss, residual temporal standardization
ORCID: https://orcid.org/0009-0007-1471-7554
相關次數: 點閱:2下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 遠端光體積變化描記法(簡稱rPPG)可由一般臉部影片估測心血管訊號,但脈搏造 成的皮膚顏色變化十分微弱,容易受到光照變化、相機反應與動作所造成的外觀變化 影響。本論文建立一套具光照穩健性的非接觸式心率估測流程,用於受控低、中、高 光照條件下的臉部影片分析。此流程先以基於PRNet的三維人臉對齊方法,將原始 影片轉換為對齊後的臉部色彩圖短片;訓練時加入短片層級的亮度與對比度增強,並 在PhysFormer 風格之時空轉換器架構的三維卷積前端之後,加入殘差式時間標準化 模組(RTSM)。模型以時間域與頻率域混合目標進行訓練:時間域項採用可容忍小幅 波形位移的軟位移皮爾森損失函數,頻率域項採用頻譜庫爾巴克-萊布勒散度損失 函數,使預測波形的頻譜分布接近由真實心率推得的目標分布,並以可調整的權重係 數𝛽控制兩者之間的平衡。
    實驗使用一個具同步臉部影片與接觸式生理訊號的實驗室遠端光體積變化描記法 資料集,並採用以受試者為單位的資料切分,其中60位受試者用於訓練,15位受試 者用於驗證。本論文評估三種實驗設定:靜態全光照混合設定、靜態與腳踏車運動全 光照混合設定,以及進一步加入說話與頭部轉動的靜態、腳踏車運動、說話及頭部轉 動混合設定。在前兩種實驗設定中,所提出的混合監督皆優於僅時間域監督與僅頻率 域監督,且𝛽=5為本研究測試範圍內表現最佳的平衡設定。貝塔權重掃描結果彙整 每種損失設定下三次獨立的完整估測器訓練,而元件消融結果則比較各個獨立的估 測器變體,其中完整估測器對應表現最佳的𝛽=5訓練結果。在靜態全光照混合設定 中,𝛽=5的三次訓練平均結果為0.99次/分鐘的心率平均絕對誤差、3.09次/分鐘的 心率均方根誤差與0.969的心率相關係數;在元件消融實驗中,完整估測器達到0.79 次/分鐘的平均絕對誤差、2.40次/分鐘的均方根誤差與0.982的心率相關係數。在靜態 與腳踏車運動全光照混合設定中,𝛽=5的三次訓練平均結果為0.95次/分鐘的平均 絕對誤差、2.53次/分鐘的均方根誤差與0.987的心率相關係數;在元件消融實驗中, 完整估測器達到0.89次/分鐘的平均絕對誤差、2.45次/分鐘的均方根誤差與0.988的 心率相關係數,相較於相同實驗設定下的PhysFormer基準模型,心率平均絕對誤差 與心率均方根誤差分別降低93.5%與86.4%。在加入說話與頭部轉動的混合設定中, 沿用前兩項實驗所選定的𝛽=5,三次報告結果的心率平均絕對誤差最小值與平均值 分別為7.65 與7.68 次/分鐘,心率均方根誤差最小值與平均值分別為16.53與16.70 次/分鐘,相關係數最大值與平均值分別為0.555與0.546。排除說話與頭部轉動短片 後,以相同三個檢查點評估靜態與腳踏車運動驗證子集,可得心率平均絕對誤差最小 值與平均值7.59與7.73次/分鐘、心率均方根誤差最小值與平均值17.19與17.47次/ 分鐘,以及相關係數最大值與平均值0.573與0.560。前兩種實驗設定的元件消融結果 顯示,所提出的混合損失函數帶來最大的單一元件改善,而短片層級光照增強與殘差 式時間標準化模組則進一步提供互補增益。殘差式時間標準化模組學得的殘差係數 在各代表性軌跡中主要呈現負值,顯示標準化後的時間特徵在模型中較像是一種抑 制性的殘差修正,可在保留原始特徵流的同時穩定時間特徵統計。這些子集結果仍差 於僅以靜態與腳踏車運動資料訓練之估測器在相同驗證子集上的結果,顯示光照穩 健性不會自動轉化為較強臉部與頭部動作下的穩健性。整體而言,本研究結果指出, 在光照變化下提升遠端光體積變化描記法之心率估測穩健性,需要同時結合輸入端 的光照多樣性、特徵層級的時間標準化,以及時間域與頻率域之間平衡的監督訊號; 對於較強的動作情境,仍需要專門的動作穩健建模方法。

    Remote photoplethysmography (rPPG) estimates cardiovascular signals from ordinary facial videos, but the pulse-induced skin-color variation is weak and easily obscured by illumina tion changes, camera response, and motion-related appearance variation. This thesis devel ops an illumination-robust camera-based heart-rate estimation pipeline for facial videos un der controlled low-, medium-, and high-illumination conditions. The pipeline uses PRNet based 3D face alignment to convert raw videos into aligned facial colormap clips, applies clip-level brightness and contrast augmentation during training, inserts a Residual Tempo ral Standardization Module (RTSM) after the 3D convolutional stem of a PhysFormer-style spatio-temporal Transformer, and trains the estimator with a hybrid temporal-frequency ob jective. ThetemporaltermisaSoft-ShiftedPearsonlossthattoleratessmallwaveformoffsets, whereas the frequency term is a spectral Kullback-Leibler divergence loss that guides the pre dicted waveform toward the ground-truth heart-rate distribution; a tuned weighting factor 𝛽 controls their balance.
    The experiments use a subject-level split (60 subjects for training and 15 subjects for val idation) of an in-lab rPPG dataset with synchronized facial videos and contact physiological signals. Three protocols are evaluated: a static all-level mix protocol, a static + bike all-level mix protocol, and a static + bike + speaking + head-rotation mix protocol that expands the evaluated motion conditions. In the first two protocols, the proposed hybrid objective out performs time-only and frequency-only supervision, and 𝛽 = 5 provides the strongest tested balance. The beta-sweep results summarize three independent full-estimator runs for each loss setting, whereas the component-ablation results compare individual estimator variants, with the full-estimator entry corresponding to the best 𝛽 = 5 run. On the static all-level mix protocol, 𝛽 = 5 achieves three-run mean performance of 0.99 bpm HR mean absolute error (MAE), 3.09 bpm HR root mean squared error (RMSE), and 0.969 HR correlation; in the component ablation study, the full estimator obtains 0.79 bpm MAE, 2.40 bpm RMSE, and 0.982 HR correlation. On the static + bike all-level mix protocol, 𝛽 = 5 achieves three-run mean performance of 0.95 bpm MAE, 2.53 bpm RMSE, and 0.987 HR correlation; in the component ablation study, the full estimator obtains 0.89 bpm MAE, 2.45 bpm RMSE, and 0.988 HR correlation, reducing HR MAE and HR RMSE by 93.5% and 86.4%, respectively, relative to the PhysFormer baseline under the sameprotocol. Forthestatic+bike+speaking+ head-rotation mix, thepreviouslyselected 𝛽 = 5settinggivesminimum/meanMAEvaluesof 7.65/7.68 bpm, minimum/mean RMSE values of 16.53/16.70 bpm, and maximum/mean cor relations of 0.555/0.546 over three reported runs. When speaking and head-rotation clips are excluded and these same three checkpoints are evaluated on the static + bikevalidation subset, they obtain minimum/mean MAE values of 7.59/7.73 bpm, minimum/mean RMSE values of 17.19/17.47 bpm, and maximum/mean correlations of 0.573/0.560. Component ablations in the first two protocols show that the proposed hybrid loss provides the largest individual im provement, while illumination augmentation and RTSM supply complementary gains. The learned RTSM residual coefficient is predominantly negative across the representative trajec tories, suggesting a suppressive residual correction that stabilizes temporal feature statistics without replacing the original feature stream. Because these results remain worse than those of the static + bike-trained estimator on the same validation subset, illumination robustness does not automatically provide robustness to stronger facial and head motion. These findings indicate that illumination-robust rPPG benefits from combining input-level illumination di versity, feature-level temporal standardization, and balanced temporal-frequency supervision, while dedicated motion-robust modeling remains necessary for more dynamic conditions.

    摘要 i Abstract iii Acknowledgment v TableofContents vi ListofTables viii ListofFigures ix ListofSymbols x 1 Introduction 1 1.1 BackgroundandMotivation 1 1.1.1 HeartRateMonitoringandPhotoplethysmography 1 1.1.2 RemotePhotoplethysmographyfromFacialVideos 2 1.1.3 IlluminationVariationasaCentralRobustnessChallenge 3 1.2 RelatedWork 4 1.2.1 EvolutionofrPPGMethods 4 1.2.2 ReportedPerformanceonCommonrPPGBenchmarks 6 1.2.3 IlluminationRobustnessandFeature-LevelNormalization 8 1.2.4 TemporalandFrequencySupervisioninrPPG 9 1.2.5 FaceAlignmentforrPPGPreprocessing 9 1.3 ProblemStatementandObjectives 11 1.3.1 ProblemStatement 11 1.3.2 ResearchObjectivesandContributions 12 1.4 Publications 12 2 Methods 13 2.1 DatasetProtocol 13 2.1.1 DataCollection 13 2.1.2 ExperimentSetup 13 2.1.3 DataSplittingandPreparation 21 2.1.4 StudyMappingandEvaluationProtocols 23 2.2 PRNet-basedPreprocessing 23 2.3 Clip-levelIlluminationAugmentation 24 2.4 Spatial-TemporalTransformerArchitecture 24 2.5 ResidualTemporalStandardizationModule 25 2.6 HybridTemporal-FrequencyLoss 26 2.7 ComputingEnvironmentandTrainingProtocol 28 2.7.1 PhysFormerBaselineProtocol 28 2.8 Heart-rateEstimationandEvaluationMetrics 29 3 Results 39 3.1 StaticAll-LevelMixStudy 39 3.1.1 ExperimentalSetting 39 3.1.2 FixedBetaSweep 39 3.1.3 PhysFormerBaselineandComponentAblation 40 3.1.4 RTSMResidualCoefficientAnalysis 40 3.2 Static+BikeAll-LevelMixStudy 41 3.2.1 ExperimentalSetting 41 3.2.2 LossContributionAnalysis 41 3.2.3 FixedBetaSweep 42 3.2.4 PhysFormerBaselineandComponentAblation 42 3.2.5 RTSMResidualCoefficientAnalysis 43 3.2.6 Illumination-LevelBreakdown 43 3.3 Static+Bike+Speaking+Head-RotationMixStudy 45 3.3.1 ExperimentalSetting 45 3.3.2 PerformanceatFixedBetaFive 45 3.3.3 RTSMResidualCoefficientAnalysis 45 4 Discussions 47 4.1 Cross-ProtocolComparison 47 4.2 RoleofHybridTemporal-FrequencySupervision 48 4.3 ConsistencyofComponentContributions 49 4.4 InterpretationofRTSM 49 4.5 ImplicationsofExtendingtheExperimentalMix 49 4.6 Validation-BasedModelSelection 50 4.7 Limitations 50 5 Conclusionsandfuturework 52 5.1 Conclusions 52 5.2 FutureWork 53 References 54 AppendixA Rightsandpermissions 59

    [1] Birla, L. and Gupta, P. (2022). AND-rPPG: A novel denoising-rPPG network for improving remote heart rate estimation. Computers in Biology and Medicine, 141:105146.
    [2] Cen, K., Fu, C.-H., and Hong, H. (2025). Robust and generalizable heart rate estimation via deep learning for remote photoplethysmography in complex scenarios. In 2025 17th International Conference on Signal Processing Systems (ICSPS), pages 687–692. IEEE.
    [3] Chen, W. and McDuff, D. (2018). Deepphys: Video-based physiological measurement using convolutional attention networks. In Proceedings of the European Conference on Computer Vision (ECCV), pages 356–373.
    [4] de Haan, G. and Jeanne, V. (2013). Robust pulse rate from chrominance-based rPPG. IEEE Transactions on Biomedical Engineering, 60(10):2878–2886.
    [5] Feng, Y., Wu, F., Shao, X., Wang, Y., and Zhou, X. (2018). Joint 3D face reconstruction and dense alignment with position map regression network. In Proceedings of the European conference on computer vision (ECCV), pages 557–574.
    [6] Hu, M., Qian, F., Guo, D., Wang, X., He, L., and Ren, F. (2021). ETA-rPPGNet: Effective time-domain attention network for remote heart rate measurement. IEEE Transactions on Instrumentation and Measurement, 70:1–12.
    [7] Huang, P.-K., Chen, T.-H., Chan, Y.-T., Chen, K.-W., and Hsu, C.-T. (2025). DD-rPPGNet: De-interfering and descriptive feature learning for unsupervised rPPG estimation. IEEE Transactions on Information Forensics and Security, 20:4956–4970.
    [8] Lee, E., Chen, E., and Lee, C.-Y. (2020). Meta-rPPG: Remote heart rate estimation using a transductive meta-learner. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 12372 LNCS:392–409.
    [9] Lee, J. S., Hwang, G., Ryu, M., and Lee, S. J. (2023). LSTC-rPPG: Long short-term convolutional network for remote photoplethysmography. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 6015–6023.
    [10] Li, Q., Guo, D., Qian, W., Tian, X., Sun, X., Zhao, H., and Wang, M. (2024). Channel-wise interactive learning for remote heart rate estimation from facial video. IEEE Transactions on Circuits and Systems for Video Technology, 34(6):4542–4555.
    [11] Li, X., Chen, J., Zhao, G., and Pietikainen, M. (2014). Remote heart rate measurement from face videos under realistic situations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4264–4271.
    [12] Li, Z. and Yin, L. (2023). Contactless pulse estimation leveraging pseudo labels and self-supervision. In Proceedings of the IEEE International Conference on Computer Vision, pages 20531–20540.
    [13] Liu, X., Fromm, J., Patel, S., and McDuff, D. (2020). Multi-task temporal shift attention networks for on-device contactless vitals measurement. Advances in Neural Information Processing Systems, 33:19400–19411.
    [14] Liu, X., Hill, B., Jiang, Z., Patel, S., and McDuff, D. (2023). EfficientPhys: Enabling simple, fast and accurate camera-based cardiac measurement. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 4997–5006.
    [15] Lu, H., Han, H., and Zhou, S. K. (2021). Dual-GAN: Joint BVP and noise modeling for remote physiological measurement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12399–12408.
    [16] Niu, X., Han, H., Shan, S., and Chen, X. (2018). Synrhythm: Learning a deep heart rate estimator from general to specific. In 2018 24th International Conference on Pattern Recognition (ICPR), pages 3580–3585. IEEE.
    [17] Niu, X., Shan, S., Han, H., and Chen, X. (2020). RhythmNet: End-to-end heart rate estimation from face via spatial-temporal representation. IEEE Transactions on Image Processing, 29:2409–2423.
    [18] Perepelkina, O., Artemyev, M., Churikova, M., and Grinenko, M. (2020). HeartTrack: Convolutional neural network for remote video-based heart rate monitoring. In IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, volume 2020-June, pages 1163–1171.
    [19] Poh, M.-Z., McDuff, D. J., and Picard, R. W. (2010). Non-contact, automated cardiac pulse measurements using video imaging and blind source separation. Optics Express, 18(10):10762–10774.
    [20] Qia, N., Li, K., Guo, D., Hu, B., and Wang, M. (2024). Cluster-phys: Facial clues clustering towards efficient remote physiological measurement. In MM 2024 - Proceedings of the 32nd ACM International Conference on Multimedia, pages 330–339.
    [21] Qian, W., Guo, D., Zhou, J., Zou, B., Yu, Z., and Wang, M. (2026). FreqPhys: Repurposing implicit physiological frequency prior for robust remote photoplethysmography. arXiv preprint arXiv:2604.00534 [cs.CV].
    [22] Shelley, K. H. and Shelley, S. L. (2001). Pulse oximeter waveform: photoelectric plethysmography. In Lake, C. L., Hines, R. L., and Blitt, C. D., editors, Clinical Monitoring: Practical Applications for Anesthesia and Critical Care, pages 420–428. W.B. Saunders.
    [23] Song, R., Chen, H., Cheng, J., Li, C., Liu, Y., and Chen, X. (2021). PulseGAN: Learning to generate realistic pulse waveforms in remote photoplethysmography. IEEE Journal of Biomedical and Health Informatics, 25(5):1373–1384.
    [24] Speth, J., Vance, N., Flynn, P., and Czajka, A. (2023). Non-contrastive unsupervised learning of physiological signals from video. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14464–14474.
    [25] Sun, Y. and Thakor, N. (2016). Photoplethysmography revisited: from contact to noncontact, from point to imaging. IEEE Transactions on Biomedical Engineering, 63(3):463–477.
    [26] Sun, Z. and Li, X. (2022). Contrast-phys: Unsupervised video-based remote physiological measurement via spatiotemporal contrast. In Computer Vision – ECCV 2022, volume 13672 of Lecture Notes in Computer Science, pages 492–510. Springer.
    [27] Tsou, Y.-Y., Lee, Y.-A., and Hsu, C.-T. (2020a). Multi-task learning for simultaneous video generation and remote photoplethysmography estimation. In Proceedings of the Asian Conference on Computer Vision, pages 392–407.
    [28] Tsou, Y.-Y., Lee, Y.-A., Hsu, C.-T., and Chang, S.-H. (2020b). Siamese-rPPG network: remote photoplethysmography signal estimation from face videos. In Proceedings of the 35th Annual ACM Symposium on Applied Computing, pages 2066–2073.
    [29] Tulyakov, S., Alameda-Pineda, X., Ricci, E., Yin, L., Cohn, J. F., and Sebe, N. (2016). Self-adaptive matrix completion for heart rate estimation from face videos under realistic conditions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2396–2404.
    [30] Verkruysse, W., Svaasand, L. O., and Nelson, J. S. (2008). Remote plethysmographic imaging using ambient light. Optics express, 16(26):21434–21445.
    [31] Špetlík, R., Franc, V., and Matas, J. (2018). Visual heart rate estimation with convolutional neural network. In Proceedings of the British Machine Vision Conference, Newcastle, UK, pages 3–6.
    [32] Wang, C.-C. (2020). Non-contact heart rate measurement based on facial videos. Master’s thesis, National Cheng Kung University, No. 1, Dasyue Rd, East District, Tainan City, 701.
    [33] Wang, K., Tang, J., Wei, Y., Liu, M., Liu, X., and Wang, Y. (2024). A plug-and-play temporal normalization module for robust remote photoplethysmography. arXiv preprint arXiv:2411.15283 [eess.IV].
    [34] Wang, W., den Brinker, A. C., Stuijk, S., and de Haan, G. (2017). Algorithmic principles of remote PPG. IEEE Transactions on Biomedical Engineering, 64(7):1479–1491.
    [35] Wang, W., Stuijk, S., and De Haan, G. (2016). A novel algorithm for remote photoplethysmography: Spatial subspace rotation. IEEE transactions on biomedical engineering, 63(9):1974–1984.
    [36] Wang, Y.-C. (2025). Comparative analysis of non-end-to-end and end-to-end deep learning models with 2D and 3D face alignment for remote heart rate estimation. Master’s thesis, National Cheng Kung University, No. 1, Dasyue Rd, East District, Tainan City, 701.
    [37] Wang, Y. C., Xu, Z. W., and Nordling, T. E. M. (2026). LSTM-based fusion of handcrafted remote photoplethysmography signals for non-contact heart rate estimation from facial video. In Proceedings of the 2026 IEEE 21st Conference on Industrial Electronics and Applications (ICIEA). IEEE. In press.
    [38] Xiong, J., Ou, W., Yao, Y., Liu, Y., Gao, Z., Liu, Z., and Gou, J. (2024). STGNet: Spatio-temporal graph neural networks considering inherent properties of physiological signals for camera-based remote photoplethysmography. Biomedical Signal Processing and Control, 98:106690.
    [39] Xu, Z. W. and Nordling, T. E. M. (2026a). Illumination-robust camera-based heart-rate estimation for physiological sensing in robots. In Proceedings of the 2026 International Conference on Advanced Robotics and Intelligent Systems (ARIS). IEEE. In press.
    [40] Xu, Z. W. and Nordling, T. E. M. (2026b). Illumination-robust camera-based heart-rate estimation for physiological sensing in robots. arXiv preprint arXiv:2606.12378 [cs.CV].
    [41] Yang, Z., Wang, H., Liu, B., and Lu, F. (2024). CbPPGGAN: A generic enhancement frame-work for unpaired pulse waveforms in camera-based photoplethysmography. IEEE Journal of Biomedical and Health Informatics, 28(2):598–608.
    [42] Yang, Z., Wang, H., and Lu, F. (2022). Assessment of deep learning-based heart rate estimation using remote photoplethysmography under different illuminations. IEEE Transactions on Human-Machine Systems, 52(6):1236–1246.
    [43] Yu, S.-N., Wang, C.-S., and Chang, Y. P. (2023a). Heart rate estimation from remote photoplethysmography based on light-weight U-Net and attention modules. IEEE Access, 11:54058–54069.
    [44] Yu, Z., Li, X., Niu, X., Shi, J., and Zhao, G. (2020). AutoHR: A strong end-to-end baseline for remote heart rate measurement with neural searching. IEEE Signal Processing Letters, 27:1245–1249.
    [45] Yu, Z., Li, X., and Zhao, G. (2019a). Remote photoplethysmograph signal measurement from facial videos using spatio-temporal networks. In Proc. BMVC.
    [46] Yu, Z., Peng, W., Li, X., Hong, X., and Zhao, G. (2019b). Remote heart rate measurement from highly compressed facial videos: an end-to-end deep learning solution with video enhancement. In Proceedings of the IEEE International Conference on Computer Vision, pages 151–160.
    [47] Yu, Z., Shen, Y., Shi, J., Zhao, H., Cui, Y., Zhang, J., Torr, P., and Zhao, G. (2023b). PhysFormer++: Facial video-based physiological measurement with SlowFast temporal difference transformer. International Journal of Computer Vision, 131(6):1307–1330. All Open Access, Hybrid Gold Open Access.
    [48] Yu, Z., Shen, Y., Shi, J., Zhao, H., Torr, P. H., and Zhao, G. (2022). PhysFormer: facial video-based physiological measurement with temporal difference transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4176–4186.
    [49] Zou, B., Zhao, Y., Hu, X., He, C., and Yang, T. (2024). Remote physiological signal recovery with efficient spatio-temporal modeling. Frontiers in Physiology, 15:1428351.

    QR CODE