PhysVR: Vision-Language Model Guided Interference-aware Temporal Feature Refinement for Remote Physiological Measurement
Abstract
Remote photoplethysmography (rPPG) enables contactless physiological measurement from facial videos, yet its subtle pulse-related variations are easily affected by illumination variation, head motion, facial blur, and region-of-interest instability. Existing methods mainly suppress interference during feature learning. However, whether interference persists in the learned temporal features, and how it can be further reduced before rPPG estimation, have received limited attention. To address this limitation, we propose PhysVR, a vision-language model guided interference-aware temporal feature refinement framework for rPPG estimation. Specifically, a physiological backbone produces global temporal features and coarse rPPG prediction, whose local temporal characteristics are used to construct time-resolved physiological reliability evidence. In parallel, a frozen vision-language model processes sampled facial frames under an interference-oriented prompt to provide visual interference evidence. Temporal cross-attention integrates the physiological and visual evidence with the global temporal features to construct interference-aware temporal context. Guided by this context, a shared temporal correction unit performs general refinement, while four interference-specific experts are adaptively routed to suppress different interference. Extensive experiments on five public benchmarks demonstrate that PhysVR consistently outperforms representative methods under both intra-dataset and cross-dataset evaluation protocols.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.