End-to-end Full Video-level Supervision for Facial Depression Assessment Model Training
Abstract
Depressive states are often reflected by long-term, variable-duration facial behaviours, yet existing video-based automatic depression assessment (ADA) approaches frequently model depression from short fixed-length segments (including frames) due to memory constraints. Segment-level supervision not only limits their ability to learn cross-segment varying-length facial behavioural patterns but also fails to capture depressive facial patterns at the full-video level. In this paper, we propose Dual-phase Memory-efficient Video-level Supervision (DMVS), a novel training strategy that equips existing segment-level face-video ADA models with end-to-end full video-level supervision through depression relevance-weighted gradient modulation. Specifically, the first phase constructs variable-length facial behaviour representations by combining features extracted from short, fixed-length segments from the entire face video, followed by estimating the depression relevance of each behaviour representation and segment-feature, based on which a video-level loss is computed from relevance-weighted behaviour representations. Then, the second phase uses each estimated depression relevance to weight the corresponding segment's contribution to the backbone gradients, guiding feature learning towards depression-informative behaviours by accumulating relevance-weighted gradients from all face video segments. This way, the target ADA model is updated based on the weighted video-level gradients. Experiments on AVEC 2013 and 2014 demonstrate that our DMVS strategy consistently improves ADA performance across ADA-specific and general-purpose spatio-temporal models, supporting our DMVS as a model-agnostic and plug-and-play training strategy. Our code is provided in Supplementary Material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.