Evidence-Oriented GRPO and Referral-Aware Screening for Mental Health
Abstract
Multimodal mental-health screening combines a high-stakes decision problem with heterogeneous audio, visual, and textual signals. Discriminative models can aggregate these signals effectively but usually expose only a score, whereas language models can consume textual evidence but need an interface that preserves the information learned by the discriminative stage. We introduce EARN (Evidence-oriented And Referral-aware screeNing using GRPO), which produces a disease probability and a 27-dimensional, symptom-named fingerprint from a shared representation. Prototype binding aligns the ordered fingerprint with a evidence library. Fingerprint verbalization renders its entries as protocolized phrases for a text-only Qwen2.5-Omni-7B reader. Evidence-oriented GRPO optimizes the discriminative evidence head separately from the reader. On the E-DAIC official split, the verbalized reader attains AUROC 0.7406, whereas the same fingerprint supplied as raw floats attains 0.5611 after fine-tuning. The proposed audio-visual pipeline obtains AUROC 0.741 and F1 0.6286 on E-DAIC, and F1 0.9527 on D-Vlog under the reported protocols. A decoupled-decision analysis further shows the trade-off between referral rate and missed positives. On E-DAIC, referring 21.4% of cases reduces automatic-channel misses from five to two. Because this analysis contains only 17 positive cases, we present it as an exploratory operating curve rather than a clinical-performance claim. Code and the evidence library are available anonymously at https://anonymous.4open.science/r/EARN-0F2D.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.