Token-wise Reward Guided Reasoning Along Diagnostic Criteria for Depression Detection
Abstract
Large Language Models (LLMs) offer broad-coverage early screening while providing natural-language descriptions of mental health risks. Fine-tuning on medical annotations is costly and risks eroding their general capabilities. Inference-time alignment offers a flexible alternative, steering frozen LLMs toward desired preferences through reward models and thereby injecting medical knowledge into depression analysis without fine-tuning the LLM. However, existing methods spread the reward signal uniformly across all tokens, diluting supervision on the few tokens that bear decisive diagnostic evidence. To address this, we propose the AlignDx framework, which recasts depression detection as a Reason-then-Assess task in which the model reasons along diagnostic dimensions grounded in established medical criteria, and a downstream classifier derives the final risk prediction from this structured analysis. During reward model training, gradient-based importance estimation quantifies each token's contribution and reweights the autoregressive reward objective accordingly, concentrating the optimization signal on decisive tokens; at inference, reward-guided decoding via logit-space fusion injects diagnostic criteria into generation. Experiments on three datasets show that AlignDx outperforms both traditional classifiers and prior LLM-based approaches in in-distribution and cross-platform settings, with gains driven by stronger high-risk detection. Our code and datasets are available at https://anonymous.4open.science/r/AlignDx-3158.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.