acceptodds
Under review as a conference paper at ICLR 2027

EviDial: When Correct Depression Predictions Lack Evidence

Abstract

Depression dialogue benchmarks typically evaluate models by comparing predicted symptom labels against clinical scale scores. Yet a correct label does not imply that the dialogue contains enough evidence to justify that prediction. We introduce EviDial, a benchmark for auditing this gap between label correctness and evidence-grounded assessment. EviDial spans 382 depression dialogues from three sources and 2,744 participant–symptom cells, characterized under an evidence taxonomy of absent, partial, conflicting, and sufficient evidence. To construct references without assuming a single gold standard, we combine generalizability theory, many-facet Rasch measurement, and random-effects latent class analysis. Our audit reveals a substantial mismatch between benchmark labels and textual evidence: 82.8% of cells lack sufficient evidence for the corresponding symptom assessment. Moreover, among label-correct predictions, 69.2–96.9% are unsupported by sufficient textual evidence across the evaluated model families and prompting strategies. We also find that high evaluator agreement can mask low measurement independence: five same-family LLM judges correspond to only 1.4 effective independent votes despite Fleiss’ κ = .785, and perturbation-based sensitivity checks fail across the tested judges. These findings show that label agreement alone can substantially overstate evidence-grounded performance in depression dialogue evaluation. EviDial provides a measurement layer for distinguishing predictions that match clinical labels from those that are justified by the evidence available in the dialogue.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.