InertiaMPDD: Multimodal Personalized Depression Detection via Individual Emotional-Inertia Modeling
Abstract
Existing omni-modal depression detection largely infers depression representations shared across individuals, rather than latent variables of individual difference inferred from each person's behavioral signals. Yet depressive expression superimposes transient states on stable individual traits: unless traits are modeled and conditioned on, they confound the assessment of states, while conditioning on annotated individual attributes does not scale because such annotations are scarce. We propose InertiaMPDD, which reframes detection as estimating each person's deviation of current expression from a counterfactual healthy baseline, turning individual differences from externally supplied labels into latent variables inferred from behavioral signals. The baseline is an individual-conditioned latent dynamics predictor trained exclusively on healthy dynamics, so departures from healthy expression systematically inflate its prediction error. Rather than regenerating a "healthy" counterpart, we realize the counterfactual by single-step rollout: under the person's own trait conditioning, the model advances the observed state one step according to healthy dynamics, and the deviation constitutes the individualized depressive signal. Because prediction and observation share the same trajectory, context and recording conditions, individual traits appear on both sides and cancel. Three complementary, annotation-free quantities follow: the residual (position), the AR(1) restoring force (speed of return to baseline), and the spectral radius of the predictor's Jacobian at the individual operating point (local stability of healthy dynamics). On the in-the-wild vlog corpora D-Vlog and LMVD, InertiaMPDD attains F1 of 98.41 and 84.98, above the best published results on both, and the joint system surpasses either pathway alone. The residual is significant with a consistent direction on both datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.