VeriRL: Grounding Video Reasoning for Depression Assessment with Verifiable Facial-Dynamics Rewards
Abstract
Reinforcement learning with verifiable rewards (RLVR) has shown strong potential for improving reasoning in multimodal large language models (MLLMs), yet extending it to medical video reasoning is challenging: task labels verify final predictions but not whether the behavioral evidence reported alongside them agrees with observable video evidence. We propose VeriRL, a post-training framework that grounds video reasoning through verifiable facial-dynamics rewards. Selected temporal facial Action Unit (AU) dynamics serve as measurable, clinically motivated behavioral references for auxiliary verification. To support progressive training, we construct two facial-evidence-aligned Chain-of-Thought corpora, GroundFace-CoT ( 4K images) and GroundVideo-CoT (1,130 video clips). During RL, the dynamics reward measures agreement between structured model reports and precomputed OpenFace-derived references. We further introduce Dynamics-Grounded Sign-GRPO (DGS-GRPO), which uses a fixed-sign classification advantage to preserve correctness feedback under homogeneous rollouts, while a bounded dynamics term adjusts its magnitude without reversing the correctness-determined sign. Experiments on AVEC 2014 and our curated Vlog dataset show gains over the supervised baseline in binary recognition, improving Accuracy/Macro-F1 from 0.61/0.61 to 0.63/0.63 on Vlog and from 0.57/0.51 to 0.60/0.55 on AVEC. On held-out Vlog videos, dynamics agreement further increases from 0.64 to 0.72, with exact channel agreement improving from 0.50 to 0.60. These results support measurable behavioral signals as auxiliary verification sources for evidence-grounded medical video reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.