SAGE: Subject-Anchored AU-Guided Expression for Pain Estimation from Video
Abstract
Pain expression is anatomically localized to a few action units (AUs) and varies strongly across individuals. We propose SAGE (Subject-Anchored AU-Guided Expression), which makes both properties explicit on top of a frozen self-supervised video transformer (MARLIN) and trains only a 2M-parameter ordinal head. AU-guided token selection routes patch tokens through the six PSPI AU regions located by facial landmarks, keeping 12% of the tokens. Subject calibration (SubjMean) centers each query on the mean of a few unlabeled clips of the same subject. On BioVid leave-one-subject-out, SAGE reaches 5-class QWK 0.399 [0.35, 0.45] (0.342 without calibration) versus 0.199 for a matched frozen baseline (PainFormer) under the same protocol. Calibration accounts for most of the gain and improves all six frozen encoders we test. It is non-parametric, and its gain rises steeply up to about five unlabeled clips per subject and only slowly beyond. Token selection improves over holistic pooling on two face backbones and outperforms both an arbitrary partition and direct AU intensities. Region attention shows that the calibrated model draws on all six AU regions, whereas the uncalibrated one collapses onto a single region. The frozen encoder transfers by linear probe to AI4Pain and zero-shot to a clinical QST pilot, exceeding PainFormer on both.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.