This study proposes a robust and cost-effective contactless remote photoplethysmography (rPPG) estimation method capable of extracting heart rate signals directly from facial videos. To address motion artifacts and environmental noise, model, designated as Balanced-TF, employs an attention-based deep learning architecture. This architecture captures long-range temporal relationships across frames while selectively attending to salient spatial features by leveraging inter-pixel relationships. We introduce a dynamic supervision strategy utilizing a hybrid loss function that incorporates both frequency and time domain losses. Time-domain supervision captures signal trends, whereas frequency-domain supervision ensures periodic physiological features are maintained within the target frequency band. The adaptive adjustment of these constraints accelerates convergence speed and minimizes overfitting risks. Extensive experimental results on the UBFC-rPPG and PURE datasets demonstrate that achieving an optimal balance between time-frequency supervision significantly enhances the robustness and generalizability of rPPG estimation. Consequently, the proposed model consistently outperforms existing methods, highlighting the efficacy of Balanced-TF in remote physiological monitoring. This study contributes to broadening the practical applicability of contactless health monitoring systems in real-world scenarios.
Sungpil Woo (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: