acceptodds
Under review as a conference paper at ICLR 2027

Token-wise Reward Guided Reasoning Along Diagnostic Criteria for Depression Detection

Abstract

Large Language Models (LLMs) offer broad-coverage early screening while providing natural-language descriptions of mental health risks. Fine-tuning on medical annotations is costly and risks eroding their general capabilities. Inference-time alignment offers a flexible alternative, steering frozen LLMs toward desired preferences through reward models and thereby injecting medical knowledge into depression analysis without fine-tuning the LLM. However, existing methods spread the reward signal uniformly across all tokens, diluting supervision on the few tokens that bear decisive diagnostic evidence. To address this, we propose the AlignDx framework, which recasts depression detection as a Reason-then-Assess task in which the model reasons along diagnostic dimensions grounded in established medical criteria, and a downstream classifier derives the final risk prediction from this structured analysis. During reward model training, gradient-based importance estimation quantifies each token's contribution and reweights the autoregressive reward objective accordingly, concentrating the optimization signal on decisive tokens; at inference, reward-guided decoding via logit-space fusion injects diagnostic criteria into generation. Experiments on three datasets show that AlignDx outperforms both traditional classifiers and prior LLM-based approaches in in-distribution and cross-platform settings, with gains driven by stronger high-risk detection. Our code and datasets are available at https://anonymous.4open.science/r/AlignDx-3158.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.