RETHINK-MH: Reliability-Triggered Multimodal Rethinking for Mental Health Screening
Abstract
Given the sensitive and identifiable nature of clinical interviews, multimodal mental health screening increasingly relies on released, de-identified measurements rather than raw recordings. Most existing systems compress these measurements into a single risk score, leaving no way to inspect which observations support a decision, which contradict it, or how conflicts between them were resolved. We present RETHINK-MH, a reliability-triggered multimodal rethinking framework that turns screening into an auditable two-pass process. A deterministic compiler organizes the acoustic, visual, and transcript measurements of each session into a two-layer evidence store: a coarse segment layer for the first-pass assessment and an indexed atomic layer that remains sealed for re-examination. An audit module estimates the error risk of each first-pass judgment from source reliability and validity-weighted cross-source disagreement, and only flagged cases trigger targeted retrieval of atomic evidence under an explicit access budget. A deterministic validator closes each case as preserved, revised, or unresolved, and unresolved or invalid cases are referred to human review, so every decision carries a checkable trail from outcome back to evidence. The loop is trained on its own cross-fitted errors via cold-start supervised fine-tuning followed by iterative preference alignment. On DAIC-WOZ, E-DAIC, and D-Vlog, RETHINK-MH achieves competitive and state-of-the-art screening performance, while correcting more first-pass errors than it introduces and recovering correct decisions for cases that would otherwise be referred to human review. Code is available at https://anonymous.4open.science/r/rethink-mh-85F3/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.