acceptodds
Under review as a conference paper at ICLR 2027

DistressBench: Evaluating Multi-Dimensional Crisis Support in Large Language Models

Abstract

Conversational language models have been adopted across everyday settings faster than the evaluation practices built to govern them. A concerning share of use is becoming affective rather than informational. Within that traffic users disclose suicide and self-harm (SSH) crises. Existing safety benchmarks nonetheless score models primarily on refusal. In this setting refusal is not the desired behavior; what the conversation requires is a substantive supportive response. We introduce DistressBench, 718 English conversations with clinician-authored user turns spanning 23 SSH subcategories, each paired with a task-specific rubric of weighted atomic criteria across seven dimensions of crisis support and a clinician reference response. A model's score is the weighted average over the criteria its reply satisfies. Because each criterion specifies a discrete action a clinician judged necessary, this quantity admits a direct operational reading as the fraction of required care delivered, rather than a position on a latent quality scale. We evaluated 25 frontier models and the best performing model satisfies 88.4% of weighted criteria and the median model 71.4%, against 98.7% for clinician references on the same instrument. However, pooled across all 25 models, in 35.3% of the conversations where a model satisfied every clinician-written criterion for recognizing the crisis, it failed to route the user toward human help. The median model leaves out nearly three in ten of the actions the conversation called for, and the user may have that conversation only once. The shortfall concentrates in the dimensions that require the model to act rather than to withhold. Contrasting each dimension's median model against its best separates capability limits from choices: de-escalation reaches 74.6% in the strongest model against a median of 46.0%, and safety disclaimers 84.6% against a median of 8.3%. Multi-turn conversations score below single-turn ones (68.5% against 74.2%). We release the corpus, the rubrics, the judge calibration labels, and a companion annex of 20 long-horizon red-team transcripts. These results argue for evaluating crisis safety on the care a model delivers rather than on what it declines to say, and for treating de-escalation and resource routing as deployment requirements.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.