簡易檢索 / 詳目顯示

研究生: 陳彥丞
CHEN, YAN CHENG
論文名稱: 基於注意力歸因與忠誠度指標之 Transformer 遠端光體積描記法可解釋性分析
Explainability Analysis of Transformer-Based Remote Photoplethysmography Using Attention Attribution and Faithfulness Metrics
指導教授: 吳馬丁
Torbjörn E. M. , Nordling
學位類別: 碩士
Master
系所名稱: 工學院 - 機械工程學系
Department of Mechanical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 253
中文關鍵詞: 擾動測試 、皮膚覆蓋率 、波形迴歸 、臉部解剖學 、稀疏路由
外文關鍵詞: perturbation testing, skin coverage, waveform regression, facial anatomy, sparse routing
相關次數: 點閱:62  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 背景。基於攝影機的遠端光體積描記術提供非接觸式心率估計途徑,近年以RhythmFormer 為代表的深度學習模型在公開基準上已達低心率誤差。然而,臨床或實務信賴仍需檢驗其視覺證據是否反映臉部具生理意義的線索,而非準確度指標無法解釋的相關性。
    方法。本論文評估四種針對 RhythmFormer 的注意力歸因方法:原始注意力、注意力展開、注意力流及 Beyond Intuition。各方法應用於 RhythmFormer 用來預測遠端光體積描記波形的稀疏注意力結構,並以兩項互補準則評估。皮膚覆蓋率衡量歸因重要性落於具解剖學相關性之未遮蔽皮膚的比例。顯著性引導忠誠度係數測試較高顯著性區域受擾動時,是否對預測波形造成較大的變化。UBFC-rPPG 為常用的公開資料集;本研究採用的公開研究群體與 RhythmFormer 實驗規程包含 42 段靜態臉部影片及同步光體積描記參考訊號。NCKU-rPPG 包含 77 位參與者的 607 個錄影時段,涵蓋不同照明條件及靜態、騎車、頭部轉動與說話四種錄影情境。
    結果。Beyond Intuition 在兩個資料集上均排名最高:在 UBFC-rPPG 上,精細皮膚覆蓋率中位數為 0.826,忠誠度係數中位數為 F = 0.917;在靜態照明第 3 級下,兩者分別為 0.789 與 0.837,但其餘方法的排名未能跨資料集轉移。注意力展開在 top-k 稀疏路由下呈現多跳洩漏,會恢復個別精細注意力層所移除的間接連結;注意力流亦有相同疑慮。將分析限定於同一情境下的單一參與者內時,兩項指標在兩個資料集上均未與單一影片片段的心率誤差、波形相關或訊噪比呈現關聯:252 個係數中有 186 個低於 |ρ| = 0.10,28 個達到 p < 0.05,而隨機情況下預期為 13 個;相較之下,將相同影片片段不區分參與者而合併分析時,名目顯著係數增至 146 個,係數幅度最高達 0.71,反映的是參與者間的差異而非影片片段間的差異。四種方法既不可互換,亦非彼此獨立:其中兩種方法依忠誠度對影片片段的排序幾乎完全相同,ρ = 0.97,但對歸因位置的一致性僅為 0.18。在八種 NCKU-rPPG 情境中,僅 Beyond Intuition 的皮膚覆蓋率與心率誤差、波形相關及訊噪比三項表現指標的變化方向均一致,分別為 ρ = −0.43、+0.57 與 +0.43;僅注意力方法的 SaCo 對三項指標皆呈相反方向。Beyond Intuition 僅在最昏暗的條件下失效:於 40 lux 時,其皮膚覆蓋率中位數降至 0.180,SaCo 中位數降至 −0.178;相較之下,動作對估計結果造成更大的劣化,卻未伴隨相同程度的下降。靜態且照明均勻的 UBFC-rPPG 呈現最低誤差,為每分鐘 0.45 次。
    結論。皮膚覆蓋率與擾動忠誠度提供互補而非可互換的證據;在兩個資料集上,兩者亦與模型表現指標互補:歸因落於皮膚不保證估計準確。歸因對一項情境所揭示的是模型關注何處,而非其歸因圖排序的忠誠程度;此結果支持 Beyond Intuition,並顯示對僅注意力方法應保持審慎。遠端光體積描記術的可解釋性評估應結合生理感知空間驗證與擾動測試,且稀疏注意力架構的歸因方法應考量其路由結構,而非僅依賴原始熱圖。

    Background. Camera-based remote photoplethysmography provides a non-contact route to heart rate estimation, and recent deep-learning models such as RhythmFormer achieve low heart-rate error on public benchmarks. However, clinical or practical trust still requires examining whether their visual evidence reflects meaningful physiological cues on the face rather than correlations that accuracy metrics cannot explain.
    Method. This thesis evaluates four attention-attribution methods for RhythmFormer: raw attention, attention rollout, attention flow, and Beyond Intuition. The methods are applied to RhythmFormer’s sparse attention structure for remote photoplethysmographic waveform prediction and assessed with two complementary criteria. Skin coverage quantifies the share of attributed importance landing on anatomically relevant uncovered skin. The Salience-guided Faithfulness Coefficient tests whether regions assigned higher salience cause larger changes in the predicted waveform when perturbed. UBFC-rPPG is a widely used public dataset; the public cohort and RhythmFormer protocol used here comprise 42 stationary facial videos with synchronized photoplethysmographic reference signals. NCKU-rPPG provided 607 sessions from 77 participants across varying illumination conditions and four recording scenarios: stationary, cycling, head rotation, and speaking.
    Results. Beyond Intuition ranked highest on both datasets, at median refined skin coverage 0.826 and F = 0.917 on UBFC-rPPG against 0.789 and 0.837 on static illumination level 3; the rankings below it did not transfer. Attention rollout exhibited multi-hop leakage under top-k sparse routing by restoring indirect connections removed from individual refined-attention layers, and attention flow shares that caution. Held within one participant of one condition, neither measure was related to the heart-rate error, the waveform correlation, or the signal-to-noise ratio of a clip on either dataset: 186 of the 252 coefficients fell below |ρ| = 0.10 and 28 reached p < 0.05 against the 13 expected by chance, whereas pooling the same clips gave 146 and reached 0.71, a measure of the participants rather than the clips. The four were neither interchangeable nor independent: two ordered the clips almost identically by faithfulness, ρ = 0.97, while agreeing on where they attributed at 0.18. Across the eight NCKU-rPPG scenarios, only Beyond Intuition’s skin coverage followed all three performance measures, at ρ = −0.43, +0.57, and +0.43; the attention-only methods’ SaCo ran opposite to each. It failed in the dimmest condition alone, its median coverage falling to 0.180 and its median SaCo to −0.178 at 40 lux, while motion degraded the estimates far more without such a drop; stationary, evenly lit UBFC-rPPG gave the lowest error, 0.45 beats per minute.
    Conclusions. Skin coverage and perturbation faithfulness are complementary rather than interchangeable, and on both datasets complementary to the performance measures: attributing to the skin does not guarantee an accurate estimate. What an attribution reveals about a condition is where the model looks rather than how faithfully its map is ordered, which supports Beyond Intuition and warrants caution with the attention-only methods. Remote photoplethysmography explainability should combine physiology-aware spatial validation with perturbation testing, while attribution for sparse-attention architectures should account for their routing structure rather than rely on raw heatmaps alone.

    Chinese abstract i Abstract iii Acknowledgment v Table of Contents vi List of Tables ix List of Figures xi List of Symbols xvi 1 Introduction 1 1.1 Remote Photoplethysmography (rPPG) 2 1.1.1 Traditional Methods 2 1.2 Deep Learning Methods for rPPG 8 1.3 Performance of rPPG Models 20 1.4 ROI Strategies and XAI Insights for deep learning rPPG 22 1.4.1 XAI 23 1.4.2 rPPG with XAI 27 1.5 Motivation 29 1.6 Problem statement and objectives 29 1.7 Contributions and thesis outline 30 2 Methods 31 2.1 Experiment Design 31 2.2 RhythmFormer Implementation Scope 31 2.2.1 Architecture and Tensor Dimensions 32 2.3 Dataset 36 2.3.1 Public Datasets 36 2.4 Data Preprocessing 37 2.4.1 Face Detection and Cropping 37 2.4.2 Temporal Clip Segmentation and Normalization 37 2.4.3 Data Split 38 2.5 Training Configuration 38 2.6 Heart Rate Estimation 39 2.6.1 Signal Post-processing 39 2.6.2 Frequency-Domain HR Estimation 39 2.7 Evaluation Metrics 40 2.8 Explainability Implementation Details 41 2.8.1 Attention Tensor Instrumentation and Extraction 42 2.8.2 Attribution Method Extensions 42 2.8.3 Beyond Intuition Extensions 45 2.8.4 SaCo Extensions 48 2.8.5 Skin-Mask Extensions 51 2.8.6 Summary of Explainability Methods 52 3 Explaining RhythmFormer 54 3.1 Introduction 54 3.2 Related Work 55 3.2.1 XAI Methods 55 3.2.2 XAI in rPPG 56 3.2.3 Quantification in XAI—Faithfulness Evaluation 56 3.3 Method 56 3.3.1 RhythmFormer Overview 56 3.3.2 Attention-Centered Attribution Methods 57 3.3.3 Faithfulness Evaluation via SaCo 60 3.3.4 Skin Attention Quantification 61 3.3.5 Waveform-Level Signal Quality 62 3.3.6 Heart-Rate-Level Error 62 3.4 Experiments 63 3.4.1 Experimental Setup 63 3.4.2 Training Reproduction 64 3.4.3 Results 65 3.5 Discussion 77 3.6 Conclusion 81 4 Cross-Dataset XAI Transfer and Model Reliability 83 4.1 Introduction 83 4.2 Related Work 84 4.3 Methods 86 4.3.1 Dataset and experimental conditions 86 4.3.2 Face extraction and resizing 88 4.3.3 Participant split and RhythmFormer preprocessing 91 4.3.4 RhythmFormer architecture 92 4.3.5 Scenario-specific training 93 4.3.6 Performance evaluation 94 4.3.7 Explainable artificial intelligence attribution methods 100 4.3.8 Anatomical plausibility and perturbation faithfulness 101 4.3.9 Cross-dataset transfer and scenario-level reliability comparisons 103 4.4 Results 104 4.4.1 Performance across illumination and motion 104 4.4.2 Attribution summaries for Static level 3 and UBFC-rPPG 106 4.4.3 Attribution summaries across the nine conditions 111 4.4.4 Scenario-level SaCo and model reliability 116 4.5 Discussion 118 4.5.1 Principal findings 118 4.5.2 Model performance across NCKU-rPPG 120 4.5.3 Static level 3 versus UBFC-rPPG 121 4.5.4 Static level 3 XAI versus UBFC-rPPG 122 4.5.5 Do skin coverage and SaCo track reliability across scenarios? 124 4.5.6 The evaluation window, and what the literature compares 125 4.5.7 Reproducibility, limitations, and future work 126 4.6 Conclusion 129 5 Discussion and Conclusions 131 5.1 Discussion 131 5.1.1 Principal findings and contributions 131 5.1.2 Quantitative XAI beyond visual inspection 132 5.1.3 Sparse routing and attribution 133 5.1.4 Cross-dataset and condition-level interpretation 134 5.1.5 Evaluation window and reference bounds 135 5.1.6 Limitations 136 5.1.7 Implications and future work 138 5.2 Data and Code Availability 139 5.3 Conclusions 139 References 141 Appendix A RhythmFormer Reproduction Details 156 A.1 Architecture and Tensor Dimensions 156 A.2 Loss and Checkpoint Selection 156 A.3 Data Augmentation 158 Appendix B XAI Feasibility Assessment of rPPG Methods 159 Appendix C Chapter 4 Supporting Details 163 C.1 Conversion contract and quality control 163 C.2 Participant and session split 164 C.3 Clip retention 164 C.4 Low-variance label checks 168 C.5 Training configuration 168 C.6 Explainable artificial intelligence results 169 C.7 Heart-rate estimation settings in the remote photoplethysmography literature since 2020 216 Appendix D Rights and Permissions 224

    [1] Abnar, S. and Zuidema, W. (2020). Quantifying attention flow in transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4190–4197. Association for Computational Linguistics.
    [2] Al-Naji, A. and Chahl, J. (2017). Contactless cardiac activity detection based on head motion magnification. International Journal of Image and Graphics, 17(01):1750001.
    [3] Álvarez Casado, C. and Bordallo López, M. (2023). Face2ppg: An unsupervised pipeline for blood volume pulse extraction from faces. IEEE Journal of Biomedical and Health Informatics, 27(11):5530–5541.
    [4] Asthana, A., Zafeiriou, S., Cheng, S., and Pantic, M. (2013). Robust discriminative response map fitting with constrained local models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3444–3451. IEEE.
    [5] Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W. (2015). On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PloS one, 10(7):e0130140.
    [6] Balakrishnan, G., Durand, F., and Guttag, J. (2013). Detecting pulse from head motions in video. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 3430–3437. IEEE.
    [7] Bhati, D., Neha, F., and Amiruzzaman, M. (2024). A survey on explainable artificial intelligence (xai) techniques for visualizing deep learning models in medical imaging. Journal of Imaging, 10(10):239.
    [8] Birla, L. and Gupta, P. (2022). AND-rPPG: A novel denoising-rPPG network for improving remote heart rate estimation. Computers in Biology and Medicine, 141:105146.
    [9] Bobbia, S., Macwan, R., Benezeth, Y., Mansouri, A., and Dubois, J. (2019). Unsupervised skin tissue segmentation for remote photoplethysmography. Pattern Recognition Letters, 124:82–90.
    [10] Bradski, G. (2000). The opencv library. Dr. Dobb’s Journal of Software Tools, 25(11):120–126.
    [11] Cao, K., Tan, T., Chen, Z., Yang, K., and Sun, Y. (2024). A novel heart rate estimation framework with self-correcting face detection for neonatal intensive care unit. Displays, 85:102852.
    [12] Castellano Ontiveros, R., Elgendi, M., and Menon, C. (2024). A machine learning-based approach for constructing remote photoplethysmogram signals from video cameras. Communications Medicine, 4(1):109.
    [13] Cen, K., Fu, C.-H., and Hong, H. (2025). Robust and generalizable heart rate estimation via deep learning for remote photoplethysmography in complex scenarios. arXiv preprint arXiv:2507.07795 [cs.CV].
    [14] Chang, J. R. and Nordling, T. E. M. (2024). Skin feature point tracking using deep feature encodings. International Journal of Machine Learning and Cybernetics, 16(4):2503–2521.
    [15] Chang, T.-Y., Wu, Y.-C., Chung, M.-L., Wu, B.-F., Cheng, H.-M., Lee, W.-C., and Chang, S.-L. (2026). Novel application of contactless remote photoplethysmography in identifying heart failure with left ventricular systolic dysfunction. Journal of Cardiology, 87(6):489–495.
    [16] Chen, J., Li, X., Yu, L., Dou, D., and Xiong, H. (2023). Beyond intuition: Rethinking token attributions inside transformers. Transactions on Machine Learning Research.
    [17] Chen, L. and Nordling, T. E. M. (2026a). Adapting RhythmFormer to a multi-condition in-lab remote photoplethysmography dataset: Preprocessing, subject-independent training, and performance across illumination and motion. [Manuscript in preparation]. Department of Mechanical Engineering, National Cheng Kung University.
    [18] Chen, L. and Nordling, T. E. M. (2026b). Explaining RhythmFormer: A systematic XAI analysis of periodic sparse attention for remote photoplethysmography. In Casalino, G., Castellano, G., Kaczmarek-Majer, K., Valerio, A. G., and Zaza, G., editors, Proceedings of the Third Workshop on Explainable Artificial Intelligence for the Medical Domain (EXPLIMED 2026), CEUR Workshop Proceedings. CEUR-WS.org. In press.
    [19] Chen, L. and Nordling, T. E. M. (2026c). Explaining RhythmFormer: A systematic XAI analysis of periodic sparse attention for remote photoplethysmography.
    [20] Chen, S., Wong, K.-L., Chin, J.-W., Chan, T.-T., and So, R. H. Y. (2024). Diffphys: Enhancing signal-to-noise ratio in remote photoplethysmography signal using a diffusion model approach. Bioengineering, 11(8):743.
    [21] Chen, W. and McDuff, D. (2018). Deepphys: Video-based physiological measurement using convolutional attention networks. In Proceedings of the European Conference on Computer Vision (ECCV), pages 356–373.
    [22] Cheng, C.-H., Wong, K.-L., Chin, J.-W., Chan, T.-T., and So, R. H. Y. (2021). Deep learning methods for remote heart rate measurement: A review and future research agenda. Sensors, 21(18):6296.
    [23] Chiu, L.-W., Chou, Y.-R., Wu, Y.-C., Chung, M.-L., Wu, B.-F., and Chou, K.-T. (2024). Video-based contactless detection of sleep apnea with deep-learning model. IEEE Transactions on Instrumentation and Measurement, 73:1–13.
    [24] Chiu, L.-W., Chou, Y.-R., Wu, Y.-C., and Wu, B.-F. (2023). Deep-learning-based remote photoplethysmography measurement in driving scenarios with color and near-infrared images. IEEE Transactions on Instrumentation and Measurement, 72:1–12.
    [25] Choi, J. and Lee, S. J. (2025). Periodic-mae: Periodic video masked autoencoder for rppg estimation.
    [26] Chu, S., Xia, M., Yuan, M., Liu, X., Seppanen, T., Zhao, G., and Shi, J. (2025a). Codephys: Robust video-based remote physiological measurement through latent codebook querying. IEEE Journal of Biomedical and Health Informatics, 29(7):4932–4945.
    [27] Chu, S., Xia, M., Yuan, M., Liu, X., Seppanen, T., Zhao, G., and Shi, J. (2025b). Codephys: Robust video-based remote physiological measurement through latent codebook querying. arXiv preprint arXiv:2502.07526 [cs.CV].
    [28] Correa, J. A. M., Abadi, M. K., Sebe, N., and Patras, I. (2021). Amigos: A dataset for affect, personality and mood research on individuals and groups. IEEE Transactions on Affective Computing, 12(2):479–493.
    [29] Dasari, A., Prakash, S. K. A., Jeni, L. s. A., and Tucker, C. S. (2021). Evaluation of biases in remote photoplethysmography methods. NPJ digital medicine, 4(1):91.
    [30] de Haan, G. and Jeanne, V. (2013). Robust pulse rate from chrominance-based rPPG. IEEE Transactions on Biomedical Engineering, 60(10):2878–2886.
    [31] De Haan, G. and Van Leest, A. (2014). Improved motion robustness of remote-ppg by using the blood volume pulse signature. Physiological measurement, 35(9):1913–1926.
    [32] De la Torre Frade, F., Chu, W.-S., Xiong, X., Carrasco, F. V., Ding, X., and Cohn, J. (2015). Intraface. In Automatic face and gesture recognition, pages 1–8.
    [33] Debnath, U. and Kim, S. (2025). A comprehensive review of heart rate measurement using remote photoplethysmography and deep learning. BioMedical Engineering OnLine, 24(1):73.
    [34] Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021. OpenReview.net.
    [35] Gao, G., Liu, L., Wang, L., and Zhang, Y. (2019). Fashion clothes matching scheme based on siamese network and autoencoder. Multimed. Syst., 25(6):593–602.
    [36] Gao, H., Zhang, C., Pei, S., and Wu, X. (2024). Region of interest analysis using delaunay triangulation for facial video-based heart rate estimation. IEEE Transactions on Instrumentation and Measurement, 73:1–12.
    [37] Guo, D., Tang, S., and Wang, M. (2019). Connectionist temporal modeling of video and language: A joint model for translation and sign labeling. In IJCAI International Joint Conference on Artificial Intelligence, volume 2019-August, pages 751–757.
    [38] Haque, M. A., Irani, R., Nasrollahi, K., and Moeslund, T. B. (2016). Heartbeat Rate Measurement from Facial Video. IEEE Intelligent Systems, 31(3):40–48.
    [39] Hassija, V., Chamola, V., Mahapatra, A., Singal, A., Goel, D., Huang, K., Scardapane, S., Spinelli, I., Mahmud, M., and Hussain, A. (2024). Interpreting black-box models: A review on explainable artificial intelligence. Cognitive Computation, 16:45–74.
    [40] Hertzman, A. B. (1937). Photoelectric plethysmography of the fingers and toes in man. Proceedings of the Society for Experimental Biology and Medicine, 37(3):529–534.
    [41] Hertzman, A. B. (1938). The blood supply of various skin areas as estimated by the photoelectric plethysmograph. American Journal of Physiology–Legacy Content, 124(2):328–340.
    [42] Heusch, G., Anjos, A., and Marcel, S. (2017). A reproducible study on remote heart rate measurement. arXiv preprint arXiv:1709.00962 [cs.CV].
    [43] Hsu, G.-S., Ambikapathi, A., and Chen, M.-S. (2017). Deep learning with time-frequency representation for pulse estimation from facial videos. In IEEE International Joint Conference on Biometrics (IJCB), volume 2018-January, pages 383–389.
    [44] Hu, M., Gao, D., Wang, X., and Tang, Y. (2024). Video segment time-channel frequency attention network for remote photoplethysmography. In 2024 5th International Conference on Computer Vision, Image and Deep Learning, CVIDL 2024, pages 1367–1371.
    [45] Hu, M., Qian, F., Guo, D., Wang, X., He, L., and Ren, F. (2021). ETA-rPPGNet: Effective time-domain attention network for remote heart rate measurement. IEEE Transactions on Instrumentation and Measurement, 70:1–12.
    [46] Hu, M., Qian, F., Wang, X., He, L., Guo, D., and Ren, F. (2022). Robust heart rate estimation with spatial-temporal attention network from facial videos. IEEE Transactions on Cognitive and Developmental Systems, 14(2):639–647.
    [47] Huang, B., Chang, C.-M., Lin, C.-L., Chen, W., Juang, C.-F., and Wu, X. (2020). Visual heart rate estimation from facial video based on cnn. In Proceedings of the 15th IEEE Conference on Industrial Electronics and Applications, ICIEA 2020, pages 1658–1662.
    [48] Huang, P.-K., Chen, T.-H., Chan, Y.-T., Chen, K.-W., and Hsu, C.-T. (2025). DD-rPPGNet: De-interfering and descriptive feature learning for unsupervised rPPG estimation. IEEE Transactions on Information Forensics and Security, 20:4956–4970.
    [49] Huang, P.-W., Wu, B.-J., and Wu, B.-F. (2021). A heart rate monitoring framework for real-world drivers using remote photoplethysmography. IEEE Journal of Biomedical and Health Informatics, 25(5):1397–1408.
    [50] Irani, R., Nasrollahi, K., and Moeslund, T. B. (2014). Improved pulse detection from head motions using dct. In 2014 international conference on computer vision theory and applications (VISAPP), volume 3, pages 118–124. IEEE.
    [51] Jain, S. and Wallace, B. C. (2019). Attention is not explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3543–3556. Association for Computational Linguistics.
    [52] Kalasampath, K., Spoorthi, K., Sajeev, S., Kuppa, S. S., Ajay, K., and Angulakshmi, M. (2025). A literature review on applications of explainable artificial intelligence (xai). IEEE Access, 13:41111–41140.
    [53] Kazemi, V. and Sullivan, J. (2014). One millisecond face alignment with an ensemble of regression trees. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, pages 1867–1874.
    [54] Knyaz, V. A., Vygolov, O., Kniaz, V. V., Vizilter, Y., Gorbatsevich, V., Luhmann, T., and Conen, N. (2017). Deep learning of convolutional auto-encoder for image matching and 3d object reconstruction in the infrared range. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pages 2155–2164.
    [55] Kolosov, D., Kelefouras, V., Kourtessis, P., and Mporas, I. (2023). Contactless camera-based heart rate and respiratory rate monitoring using ai on hardware. Sensors, 23(9):4550. All Open Access, Gold Open Access, Green Open Access.
    [56] Kuang, H., Lv, F., Ma, X., and Liu, X. (2022). Efficient spatiotemporal attention network for remote heart rate variability analysis. Sensors, 22(3):1010. All Open Access, Gold Open Access, Green Open Access.
    [57] Kwon, S., Kim, J., Lee, D., and Park, K. (2015). Roi analysis for remote photoplethysmography on facial video. In 2015 37th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), pages 4938–4941.
    [58] Lee, E., Chen, E., and Lee, C.-Y. (2020). Meta-rPPG: Remote heart rate estimation using a transductive meta-learner. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 12372 LNCS:392–409.
    [59] Lee, J. S., Hwang, G., Ryu, M., and Lee, S. J. (2023a). LSTC-rPPG: Long short-term convolutional network for remote photoplethysmography. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 6015–6023.
    [60] Lee, R. J., Sivakumar, S., and Lim, K. H. (2024). Review on remote heart rate measurements using photoplethysmography. Multimedia Tools and Applications, 83(15):44699–44728.
    [61] Lee, S., Lee, M., and Sim, J. Y. (2023b). Dse-nn: Deeply supervised efficient neural network for real-time remote photoplethysmography. Bioengineering, 10(12):1428. All Open Access, Gold Open Access, Green Open Access.
    [62] Lewandowska, M., Rumiński, J., Kocejko, T., and Nowak, J. (2011). Measuring pulse rate with a webcam—a non-contact method for evaluating cardiac activity. In 2011 federated conference on computer science and information systems (FedCSIS), pages 405–410. IEEE.
    [63] Li, H., Lu, H., and Chen, Y.-C. (2024a). Bi-tta: Bidirectional test-time adapter for remote physiological measurement. In Computer Vision – ECCV 2024, volume 15069 of Lecture Notes in Computer Science, pages 356–374. Springer.
    [64] Li, J., Yu, Z., and Shi, J. (2023). Learning motion-robust remote photoplethysmography through arbitrary resolution videos. Proceedings of the AAAI Conference on Artificial Intelligence, 37(1):1334–1342.
    [65] Li, K., Guo, D., and Wang, M. (2021). Proposal-free video grounding with contextual pyramid network. In 35th AAAI Conference on Artificial Intelligence, AAAI 2021, volume 35, pages 1902–1910. All Open Access, Gold Open Access.
    [66] Li, Q., Guo, D., Qian, W., Tian, X., Sun, X., Zhao, H., and Wang, M. (2024b). Channel-wise interactive learning for remote heart rate estimation from facial video. IEEE Transactions on Circuits and Systems for Video Technology, 34(6):4542–4555.
    [67] Li, X., Alikhani, I., Shi, J., Seppanen, T., Junttila, J., Majamaa-Voltti, K., Tulppo, M., and Zhao, G. (2018). The obf database: a large face video database for remote physiological signal measurement and atrial fibrillation detection. In 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018), pages 242–249. IEEE.
    [68] Li, X., Chen, J., Zhao, G., and Pietikainen, M. (2014). Remote heart rate measurement from face videos under realistic situations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4264–4271.
    [69] Li, Y., Huang, J., Zhao, J., Wu, D., and Zheng, M. (2025). Ts-can+: A improved ts-can architecture for non-contact heart rate measurement. IEEE Transactions on Consumer Electronics, 71(1):1393–1401.
    [70] Li, Z., Wu, X., Álvarez Casado, C., Lindholm, V., Mikkonen, K., Xia, Z., Feng, X., and Bordallo López, M. (2026). A comprehensive survey on contactless vital sign monitoring using vision-based, radio-based, and fusion approaches. Neurocomputing, 674:132877.
    [71] Li, Z. and Yin, L. (2023). Contactless pulse estimation leveraging pseudo labels and self-supervision. In Proceedings of the IEEE International Conference on Computer Vision, pages 20531–20540.
    [72] Lin, H.-C., Dong, Y.-C., Wu, B.-J., and Wu, B.-F. (2022). Measurement of blood oxygen based on remote-photoplethysmography. In 2022 International Conference on System Science and Engineering (ICSSE), pages 121–126. IEEE.
    [73] Liu, I., Liu, F., Zhong, Q., Ma, F., and Ni, S. (2024a). Your blush gives you away: detecting hidden mental states with remote photoplethysmography and thermal imaging. PeerJ Computer Science, 10:e1912.
    [74] Liu, I., Liu, F., Zhong, Q., Ma, F., and Ni, S. (2024b). Your blush gives you away: detecting hidden mental states with remote photoplethysmography and thermal imaging.
    [75] Liu, J., Zhang, J., Tang, T., and Wu, S. (2025a). Dadnet: text detection of arbitrary shapes from drone perspective based on boundary adaptation. Complex and Intelligent Systems, 11(1).
    [76] Liu, S.-Q. and Yuen, P. C. (2020). A general remote photoplethysmography estimator with spatiotemporal convolutional network. In 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020), pages 481–488. IEEE.
    [77] Liu, T., Xiao, H., Sun, Y., Zuo, K., Li, Z., Yang, Z., and Liu, S. (2025b). Physkannet: A kan-based model for multiscale feature extraction and contextual fusion in remote physiological measurement. Biomedical Signal Processing and Control, 100:107111.
    [78] Liu, X., Fromm, J., Patel, S., and McDuff, D. (2020). Multi-task temporal shift attention networks for on-device contactless vitals measurement. Advances in Neural Information Processing Systems, 33:19400–19411.
    [79] Liu, X., Hill, B., Jiang, Z., Patel, S., and McDuff, D. (2023a). EfficientPhys: Enabling simple, fast and accurate camera-based cardiac measurement. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 4997–5006.
    [80] Liu, X., McDuff, D., Narayanswamy, G., Paruchuri, A., Patel, S., Sengupta, R., Tang, J., and Wang, Y. (2023b). rppg-toolbox: Deep remote ppg toolbox. In Advances in Neural Information Processing Systems 36, pages 68485–68510.
    [81] Liu, X., Wei, W., Kuang, H., and Ma, X. (2022). Heart rate measurement based on 3D central difference convolution with attention mechanism. Sensors, 22(2):688.
    [82] Lu, H., Han, H., and Zhou, S. K. (2021). Dual-GAN: Joint BVP and noise modeling for remote physiological measurement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12399–12408.
    [83] Lundberg, S. M. and Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30, pages 4765–4774.
    [84] Matthes, K. and Hauss, W. (1938). Lichtelektrische plethysmogramme. Klinische Wochenschrift, 17(35):1211–1213.
    [85] McDuff, D., Wander, M., Liu, X., Hill, B., Hernandez, J., Lester, J., and Baltrusaitis, T. (2022). Scamps: Synthetics for camera measurement of physiological signals. Advances in Neural Information Processing Systems, 35:3744–3757.
    [86] Mersha, M., Lam, K., Wood, J., AlShami, A. K., and Kalita, J. (2024). Explainable artificial intelligence: A survey of needs, techniques, applications, and future direction. Neurocomputing, 599:128111.
    [87] Mounsey, P. (1957). Praecordial ballistocardiography. British heart journal, 19(2):259–271.
    [88] Nakajima, K., Maekawa, T., and Miike, H. (1997). Detection of apparent skin motion using optical flow analysis: Blood pulsation signal obtained from optical flow sequence. Review of scientific instruments, 68(2):1331–1336.
    [89] Narayan, K., VS, V., Chellappa, R., and Patel, V. M. (2025). FaceXFormer: A unified transformer for facial analysis. In 2025 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11369–11382. IEEE.
    [90] Nguyen, N., Nguyen, L., Álvarez Casado, C., Silvén, O., and Bordallo López, M. (2023). Non-contact heart rate measurement from deteriorated videos. In 2023 IEEE 28th International Conference on Emerging Technologies and Factory Automation (ETFA), pages 1–8. IEEE.
    [91] Ni, A., Azarang, A., and Kehtarnavaz, N. (2021). A review of deep learning-based contactless heart rate measurement methods. Sensors, 21(11):3719.
    [92] Niu, X., Han, H., Shan, S., and Chen, X. (2018a). Synrhythm: Learning a deep heart rate estimator from general to specific. In 2018 24th International Conference on Pattern Recognition (ICPR), pages 3580–3585. IEEE.
    [93] Niu, X., Han, H., Shan, S., and Chen, X. (2018b). Vipl-hr: A multi-modal database for pulse estimation from less-constrained face video. In Asian Conference on Computer Vision, pages 562–576. Springer.
    [94] Niu, X., Shan, S., Han, H., and Chen, X. (2020a). RhythmNet: End-to-end heart rate estimation from face via spatial-temporal representation. IEEE Transactions on Image Processing, 29:2409–2423.
    [95] Niu, X., Yu, Z., Han, H., Li, X., Shan, S., and Zhao, G. (2020b). Video-based remote physiological measurement via cross-verified feature disentangling. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK , August 23–28, 2020, Proceedings, Part II 16, volume 12347 LNCS of Lecture Notes in Computer Science, pages 295–310. Springer.
    [96] Niu, X., Zhao, X., Han, H., Das, A., Dantcheva, A., Shan, S., and Chen, X. (2019). Robust remote heart rate estimation from face utilizing spatial-temporal attention. In Proceedings - 14th IEEE International Conference on Automatic Face and Gesture Recognition, FG 2019, pages 1–8. All Open Access, Green Open Access.
    [97] Perepelkina, O., Artemyev, M., Churikova, M., and Grinenko, M. (2020). HeartTrack: Convolutional neural network for remote video-based heart rate monitoring. In IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, volume 2020-June, pages 1163–1171.
    [98] Petsiuk, V., Das, A., and Saenko, K. (2018). RISE: Randomized input sampling for explanation of black-box models. In Proceedings of the British Machine Vision Conference (BMVC).
    [99] Pilz, C. S., Blazek, V., and Leonhardt, S. (2019). On the vector space in photoplethysmography imaging. In 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), pages 1580–1588. IEEE.
    [100] Pilz, C. S., Krajewski, J., and Blazek, V. (2017). On the diffusion process for heart rate estimation from face videos under realistic conditions. In German Conference on Pattern Recognition, pages 361–373. Springer.
    [101] Pilz, C. S., Zaunseder, S., Krajewski, J., and Blazek, V. (2018). Local group invariance for heart rate estimation from face videos in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 1254–1262.
    [102] Poh, M.-Z., Kim, K., Goessling, A., Swenson, N., and Picard, R. (2012). Cardiovascular monitoring using earphones and a mobile device. IEEE Pervasive Computing, 11(4):18–26.
    [103] Pruthi, D., Gupta, M., Dhingra, B., Neubig, G., and Lipton, Z. C. (2020). Learning to deceive with attention-based explanations. In Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J., editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4782–4793, Online. Association for Computational Linguistics.
    [104] PyTorch Contributors (2024). Reproducibility. https://pytorch.org/docs/stable/notes/randomness.html. Accessed 2026-05-07.
    [105] Qi, H., Guo, Z., Chen, X., Shen, Z., and Wang, Z. J. (2017). Video-based human heart rate measurement using joint blind source separation. Biomedical Signal Processing and Control, 31:309–320.
    [106] Qia, N., Li, K., Guo, D., Hu, B., and Wang, M. (2024). Cluster-phys: Facial clues clustering towards efficient remote physiological measurement. In MM 2024 - Proceedings of the 32nd ACM International Conference on Multimedia, pages 330–339.
    [107] Qian, W., Guo, D., Li, K., Zhang, X., Tian, X., Yang, X., and Wang, M. (2024). Dual-path tokenlearner for remote photoplethysmography-based physiological measurement with facial videos. IEEE Transactions on Computational Social Systems, 11(3):4465–4477.
    [108] Qian, W., Guo, D., Zhou, J., Zou, B., Yu, Z., and Wang, M. (2026). FreqPhys: Repurposing implicit physiological frequency prior for robust remote photoplethysmography. arXiv preprint arXiv:2604.00534 [cs.CV].
    [109] Qiu, Y., Liu, Y., Arteaga-Falconi, J., Dong, H., and Saddik, A. E. (2019). Evm-cnn: Real-time contactless heart rate estimation from facial video. IEEE Transactions on Multimedia, 21(7):1778–1787. All Open Access, Green Open Access.
    [110] Ren, W., Chen, Y., Zhang, D., and Karimi, H. R. (2024). Bi-level weighted mixed-domain self-attention network for non-contact heart rate estimation. Knowledge-Based Systems, 300:112262.
    [111] Ribeiro, M. T., Singh, S., and Guestrin, C. (2016). "why should i trust you?": Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1135–1144.
    [112] Sahoo, N. N., Sachidanand, V., Gayathri, M. N., Murugesan, B., Ram, K., Joseph, J., and Sivaprakasam, M. (2025). Kdphys: An attention guided 3d to 2d knowledge distillation for real-time video-based physiological measurement. Biomedical Signal Processing and Control, 107:107797.
    [113] Sahoo, N. N., Sachidanand, V., Gayathri, M. N., Murugesan, B., Ram, K., Joseph, J., and Sivaprakasam, M. (2026). Kdphys: An attention guided 3d to 2d knowledge distillation for real-time video-based physiological measurement. arXiv preprint arXiv:2601.00714 [eess.IV].
    [114] Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2019). Grad-cam: Visual explanations from deep networks via gradient-based localization. International Journal of Computer Vision, 128(2):336–359.
    [115] Serrano, S. and Smith, N. A. (2019). Is attention interpretable? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2931–2951. Association for Computational Linguistics.
    [116] Shan, L. and Yu, M. (2013). Video-based heart rate measurement using head motion tracking and ica. In 2013 6th International Congress on Image and Signal Processing (CISP), volume 1, pages 160–164. IEEE.
    [117] Simonyan, K., Vedaldi, A., and Zisserman, A. (2014). Deep inside convolutional networks: Visualising image classification models and saliency maps. In 2nd International Conference on Learning Representations, ICLR 2014, Workshop Track Proceedings.
    [118] Smilkov, D., Thorat, N., Kim, B., Viégas, F., and Wattenberg, M. (2017). SmoothGrad: removing noise by adding noise. arXiv preprint arXiv:1706.03825 [cs.LG].
    [119] Soleymani, M., Lichtenauer, J., Pun, T., and Pantic, M. (2012). A multimodal database for affect recognition and implicit tagging. IEEE Transactions on Affective Computing, 3(1):42–55.
    [120] Song, R., Chen, H., Cheng, J., Li, C., Liu, Y., and Chen, X. (2021). PulseGAN: Learning to generate realistic pulse waveforms in remote photoplethysmography. IEEE Journal of Biomedical and Health Informatics, 25(5):1373–1384.
    [121] Speth, J., Vance, N., Flynn, P., and Czajka, A. (2023). Non-contrastive unsupervised learning of physiological signals from video. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14464–14474.
    [122] Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M. (2015). Striving for simplicity: The all convolutional net. In ICLR Workshop.
    [123] Stricker, R., Müller, S., and Groß, H.-M. (2014). Non-contact video-based pulse rate measurement on a mobile service robot. In The 23rd IEEE International Symposium on Robot and Human Interactive Communication, pages 1056–1062, Edinburgh, Scotland, UK. IEEE, IEEE. PURE dataset.
    [124] Sun, N., Wang, N., Liu, J., Chai, L., and Sun, H. (2025). Remote heart rate measurement based on spatial–temporal self-attention. Computers and Electrical Engineering, 127:110557.
    [125] Sun, Y., Yang, Y.-Y., Wu, B.-J., Huang, P.-W., Cheng, S.-E., Wu, B.-F., and Chen, C.-C. (2022). Contactless facial video recording with deep learning models for the detection of atrial fibrillation. Scientific Reports, 12(1):281.
    [126] Sun, Z. and Li, X. (2022). Contrast-phys: Unsupervised video-based remote physiological measurement via spatiotemporal contrast. In Computer Vision – ECCV 2022, volume 13672 of Lecture Notes in Computer Science, pages 492–510. Springer.
    [127] Sun, Z. and Li, X. (2024). Contrast-phys+: Unsupervised and weakly-supervised video-based remote physiological measurement via spatiotemporal contrast. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(8):5835–5851.
    [128] Sundararajan, M., Taly, A., and Yan, Q. (2017). Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 3319–3328. PMLR.
    [129] Taiwan Food and Drug Administration (2020). Technical guidance for premarket registration of artificial intelligence/machine learning-based software as a medical device. https://www.fda.gov.tw/tc/includes/GetFile.ashx?id=f637354440708445688&type=1. Accessed 2026-05-17.
    [130] Tang, J., Chen, K., Wang, Y., Shi, Y., Patel, S., McDuff, D., and Liu, X. (2023). Mmpd: multi-domain mobile video physiology dataset. In 2023 45th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), pages 1–5. IEEE.
    [131] Tseng, C.-W., Wu, B.-F., and Sun, Y. (2025). A real-time contact-free atrial fibrillation detection system for mobile devices. IEEE Journal of Biomedical and Health Informatics, 29(1):17–29.
    [132] Tsou, Y.-Y., Lee, Y.-A., and Hsu, C.-T. (2020a). Multi-task learning for simultaneous video generation and remote photoplethysmography estimation. In Proceedings of the Asian Conference on Computer Vision, pages 392–407.
    [133] Tsou, Y.-Y., Lee, Y.-A., Hsu, C.-T., and Chang, S.-H. (2020b). Siamese-rPPG network: remote photoplethysmography signal estimation from face videos. In Proceedings of the 35th Annual ACM Symposium on Applied Computing, pages 2066–2073.
    [134] Tu, X., Liu, Y., Zhao, M., Liu, B., Liu, J., Lei, X., Xu, L., Zhu, X., Wang, Y., and Huang, Y. (2024). Standardized rppg signal generation based on generative adversarial networks. Journal of Electronic Imaging, 33(1).
    [135] Tulyakov, S., Alameda-Pineda, X., Ricci, E., Yin, L., Cohn, J. F., and Sebe, N. (2016). Self-adaptive matrix completion for heart rate estimation from face videos under realistic conditions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2396–2404.
    [136] U.S. Food and Drug Administration, Health Canada, and U.K. Medicines and Healthcare products Regulatory Agency (2024). Transparency for machine learning-enabled medical devices: Guiding principles. https://www.fda.gov/media/179269/download. Accessed 2026-05-17.
    [137] van Putten, L. D., Bamford, K. E., Veleslavov, I., and Wegerif, S. (2024). From video to vital signs: using personal device cameras to measure pulse rate and predict blood pressure using explainable ai. Discover Applied Sciences, 6(4):184.
    [138] Verkruysse, W., Svaasand, L. O., and Nelson, J. S. (2008). Remote plethysmographic imaging using ambient light. Optics express, 16(26):21434–21445.
    [139] Viola, P. and Jones, M. J. (2004). Robust real-time face detection. International journal of computer vision, 57(2):137–154.
    [140] Špetlı́k, R., Franc, V., and Matas, J. (2018). Visual heart rate estimation with convolutional neural network. In Proceedings of the British Machine Vision Conference, Newcastle, UK, pages 3–6.
    [141] Wachter, S., Mittelstadt, B., and Russell, C. (2018). Counterfactual explanations without opening the black box: Automated decisions and the gdpr. Harvard Journal of Law & Technology, 31(2):841–887.
    [142] Wang, C.-C. (2020). Non-contact heart rate measurement based on facial videos. Master’s thesis, National Cheng Kung University, No. 1, Dasyue Rd, East District, Tainan City, 701.
    [143] Wang, K., Tang, J., Wei, Y., Liu, M., Liu, X., and Wang, Y. (2024). A plug-and-play temporal normalization module for robust remote photoplethysmography. arXiv preprint arXiv:2411.15283 [eess.IV].
    [144] Wang, W., Den Brinker, A. C., and De Haan, G. (2019). Single-element remote-ppg. IEEE Transactions on Biomedical Engineering, 66(7):2032–2043.
    [145] Wang, W., Den Brinker, A. C., and De Haan, G. (2020). Discriminative signatures for remote-ppg. IEEE Transactions on Biomedical Engineering, 67(5):1462–1473.
    [146] Wang, W., den Brinker, A. C., Stuijk, S., and de Haan, G. (2017). Algorithmic principles of remote PPG. IEEE Transactions on Biomedical Engineering, 64(7):1479–1491.
    [147] Wang, W., Stuijk, S., and De Haan, G. (2015). Exploiting spatial redundancy of image sensor for motion robust rPPG . IEEE Transactions on Biomedical Engineering, 62(2):415–425.
    [148] Wang, W., Stuijk, S., and De Haan, G. (2016). A novel algorithm for remote photoplethysmography: Spatial subspace rotation. IEEE transactions on biomedical engineering, 63(9):1974–1984.
    [149] Wang, Y.-C. (2025). Comparative analysis of non-end-to-end and end-to-end deep learning models with 2D and 3D face alignment for remote heart rate estimation. Master’s thesis, National Cheng Kung University, No. 1, Dasyue Rd, East District, Tainan City, 701.
    [150] Wang, Y. C., Xu, Z. W., and Nordling, T. E. M. (2026). LSTM-based fusion of handcrafted remote photoplethysmography signals for non-contact heart rate estimation from facial video. In Proceedings of the 2026 IEEE 21st Conference on Industrial Electronics and Applications (ICIEA). IEEE. In press.
    [151] Wiegreffe, S. and Pinter, Y. (2019). Attention is not not explanation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 11–20. Association for Computational Linguistics.
    [152] Wu, B.-F., Chu, Y.-W., Huang, P.-W., and Chung, M.-L. (2019a). Neural network based luminance variation resistant remote-photoplethysmography for driver’s heart rate monitoring. IEEE Access, 7:57210–57225.
    [153] Wu, B.-F., Huang, P.-W., He, D.-H., Lin, C.-H., and Chen, K.-H. (2019b). Remote photoplethysmography enhancement with machine leaning methods. In 2019 IEEE International Conference on Systems, Man and Cybernetics (SMC), pages 2466–2471. IEEE.
    [154] Wu, B.-F., Huang, P.-W., Lin, C.-H., Chung, M.-L., Tsou, T.-Y., and Wu, Y.-L. (2018). Motion resistant image-photoplethysmography based on spectral peak tracking algorithm. IEEE Access, 6:21621–21634.
    [155] Wu, B.-F., Wu, B.-J., Cheng, S.-E., Sun, Y., and Chung, M.-L. (2023a). Motion-robust atrial fibrillation detection based on remote-photoplethysmography. IEEE Journal of Biomedical and Health Informatics, 27(6):2705–2716.
    [156] Wu, B.-F., Wu, Y.-C., and Chou, Y.-W. (2022). A compensation network with error mapping for robust remote photoplethysmography in noise-heavy conditions. IEEE Transactions on Instrumentation and Measurement, 71:1–11.
    [157] Wu, B.-F., Yang, Y.-Y., Tsai, B.-R., Huang, P.-W., Tsai, Y.-C., and Chen, K.-H. (2019c). Remote heartrate measurement based on signal feature detection in time domain. In 2019 International Conference on System Science and Engineering (ICSSE), pages 88–93. IEEE.
    [158] Wu, J., Kang, W., Tang, H., Hong, Y., and Yan, Y. (2024a). On the faithfulness of vision transformer explanations. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10936–10945.
    [159] Wu, Y.-C., Chiu, L.-W., Wu, B.-F., Lin, L. L.-C., Ho, T.-H., Chung, M.-L., and Wu, S.-F. (2023b). Motion robust remote photoplethysmography measurement during exercise for contactless physical activity intensity detection. IEEE Transactions on Instrumentation and Measurement, 72:1–14.
    [160] Wu, Y.-C., Lin, C.-H., Chiu, L.-W., Wu, B.-F., Chung, M.-L., Tang, S.-C., and Sun, Y. (2024b). Contact-free atrial fibrillation screening with attention network. IEEE Journal of Biomedical and Health Informatics, 28(9):5124–5135.
    [161] Xiang, J. and Zhu, G. (2017). Joint face detection and facial expression recognition with mtcnn. In 2017 4th international conference on information science and control engineering (ICISCE), pages 424–427. IEEE.
    [162] Xiao, H., Li, Z., Xia, Z., Liu, T., Zhou, F., and Avolio, A. (2024). Simfupulse: A self-similarity supervised model for remote photoplethysmography extraction from facial videos. Biomedical Signal Processing and Control, 98:106736.
    [163] Xiong, J., Ou, W., Yao, Y., Liu, Y., Gao, Z., Liu, Z., and Gou, J. (2024). STGNet: Spatio-temporal graph neural networks considering inherent properties of physiological signals for camera-based remote photoplethysmography. Biomedical Signal Processing and Control, 98:106690.
    [164] Xiong, X. and De la Torre, F. (2013). Supervised descent method and its applications to face alignment. In 2013 IEEE Conference on Computer Vision and Pattern Recognition, pages 532–539. IEEE.
    [165] Yang, Z., Wang, H., Liu, B., and Lu, F. (2024). CbPPGGAN: A generic enhancement framework for unpaired pulse waveforms in camera-based photoplethysmography. IEEE Journal of Biomedical and Health Informatics, 28(2):598–608.
    [166] Yang, Z., Wang, H., and Lu, F. (2022). Assessment of deep learning-based heart rate estimation using remote photoplethysmography under different illuminations. IEEE Transactions on Human-Machine Systems, 52(6):1236–1246.
    [167] Yu, C., Wang, J., Peng, C., Gao, C., Yu, G., and Sang, N. (2018). BiSeNet: Bilateral segmentation network for real-time semantic segmentation. In Computer Vision – ECCV 2018, pages 334–349.
    [168] Yu, S.-N., Wang, C.-S., and Chang, Y. P. (2023a). Heart rate estimation from remote photoplethysmography based on light-weight U-Net and attention modules. IEEE Access, 11:54058–54069.
    [169] Yu, Z., Li, X., Niu, X., Shi, J., and Zhao, G. (2020). AutoHR: A strong end-to-end baseline for remote heart rate measurement with neural searching. IEEE Signal Processing Letters, 27:1245–1249.
    [170] Yu, Z., Li, X., and Zhao, G. (2019a). Remote photoplethysmograph signal measurement from facial videos using spatio-temporal networks. In Proc. BMVC.
    [171] Yu, Z., Li, X., and Zhao, G. (2021). Facial-video-based physiological signal measurement: Recent advances and affective applications. IEEE Signal Processing Magazine, 38(6):50–58.
    [172] Yu, Z., Peng, W., Li, X., Hong, X., and Zhao, G. (2019b). Remote heart rate measurement from highly compressed facial videos: an end-to-end deep learning solution with video enhancement. In Proceedings of the IEEE International Conference on Computer Vision, pages 151–160.
    [173] Yu, Z., Shen, Y., Shi, J., Zhao, H., Cui, Y., Zhang, J., Torr, P., and Zhao, G. (2023b). PhysFormer++: Facial video-based physiological measurement with SlowFast temporal difference transformer. International Journal of Computer Vision, 131(6):1307–1330. All Open Access, Hybrid Gold Open Access.
    [174] Yu, Z., Shen, Y., Shi, J., Zhao, H., Cui, Y., Zhang, J., Torr, P., and Zhao, G. (2023c). Physformer++: Facial video-based physiological measurement with slowfast temporal difference transformer. arXiv preprint arXiv:2302.03548 [cs.CV].
    [175] Yu, Z., Shen, Y., Shi, J., Zhao, H., Torr, P. H., and Zhao, G. (2022). PhysFormer: facial video-based physiological measurement with temporal difference transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4176–4186.
    [176] Zeiler, M. D. and Fergus, R. (2014). Visualizing and understanding convolutional networks. In Computer Vision – ECCV 2014, pages 818–833.
    [177] Zhang, J., Sun, H., Hu, Y., Zhu, G., Liu, F., Yan, B., Pu, J., Du, X., Liu, J., Liu, L., et al. (2025). Channel attention pyramid network for remote physiological measurement. Scientific Reports, 15(1):22495.
    [178] Zhang, K., Zhang, Z., Li, Z., and Qiao, Y. (2016). Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Processing Letters, 23(10):1499–1503.
    [179] Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A. (2016). Learning deep features for discriminative localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2921–2929.
    [180] Zhou, K., Krause, S., Blocher, T., and Stork, W. (2020). Enhancing remote-ppg pulse extraction in disturbance scenarios utilizing spectral characteristics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 1130–1138.
    [181] Zhu, L., Wang, X., Ke, Z., Zhang, W., and Lau, R. W. H. (2023). BiFormer: Vision transformer with bi-level routing attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10323–10333.
    [182] Zou, B., Guo, Z., Chen, J., Zhuo, J., Huang, W., and Ma, H. (2024a). Rhythmformer: Extracting patterned rppg signals based on periodic sparse attention. arXiv preprint arXiv:2402.12788 [cs.CV].
    [183] Zou, B., Guo, Z., Chen, J., Zhuo, J., Huang, W., and Ma, H. (2025). Rhythmformer: Extracting patterned rPPG signals based on periodic sparse attention. Pattern Recognition, 164:111511.
    [184] Zou, B., Zhao, Y., Hu, X., He, C., and Yang, T. (2024b). Remote physiological signal recovery with efficient spatio-temporal modeling. Frontiers in Physiology, 15:1428351.

    下載圖示
    校外:立即公開
    QR CODE