Risk-Calibrated Segment-Aware Temporal Evidence Graphs for Multimodal Depression Screening
Abstract
abstract Multimodal depression screening aims to assess depression risk using behavioral cues observed during clinical interviews, such as speech, facial movements, and head pose. Existing methods often compress an entire interview into a single global representation, potentially weakening stage-dependent evidence, and require a binary decision for every subject, even when uncertainty is high. To address these limitations, we propose a segment-aware temporal evidence graph with selective risk control. Our method treats each window–modality pair as a graph node and combines temporal encoding with typed graph attention to capture temporal dependencies and cross-modal interactions. Segment-aware pooling then aggregates global, early-stage, middle-stage, and late-stage representations, along with the difference between late-stage and early-stage representations, for subject-level prediction. To handle uncertain cases, we introduce a positive-priority, two-round selective screening strategy that combines probability calibration, asymmetric decision thresholds, and ensemble disagreement. Ablation studies on the DAIC-WOZ development set support the effectiveness of segment-aware pooling and the complementary contributions of audio and visual behavioral cues. Selective screening results further show that the proposed strategy explicitly balances automated prediction coverage and the risk of missing positive cases by abstaining from predictions for uncertain subjects. Code is available at: https://anonymous.4open.science/r/rc-sateg-anonymous-3876/. abstract
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.