acceptodds
Under review as a conference paper at ICLR 2027

SafeDrift: Reliability-Audited Longitudinal Safety Evaluation of Memory-Augmented Mental-Health Agents

Abstract

Longitudinal safety evaluation of memory-augmented mental-health companion agents stacks several measurement choices (which rater scores a response, how dimension-level verdicts are aggregated, and how failures are attributed to memory) on top of the agent behavior it aims to measure. We introduce SafeDrift, an evaluation framework in which 48 clinician-reviewed personas converse over 10–15 sessions with three open-weight subject models under three memory conditions (none, rolling summary, retrieval), with every safety-critical user probe delivered verbatim-identically across all nine model–memory cells (preceding conversational histories still differ). We audit the measurement pipeline itself: a stratified 731-verdict sample is scored by the benchmark's judge, two reference LLM judges, and two human annotators. Reliability is strongly dimension-specific: referral and sycophancy fall below 0.6 under both chance-corrected agreement measures, Cohen's κ and Gwet's AC1; the annotators disagree with each other on referral (κ = 0.30), with their calibration responses indicating two distinct interpretations of the rubric; and both annotators are systematically stricter than every LLM judge on sycophancy. We then ask which substantive conclusions survive. The recognition-failure gap between the smallest and the two larger subject models (≥2.6×) replicates under all five raters; the analogous judge-derived referral gap does not. Pooled over models, memory-condition contrasts are near zero: five of six satisfy a ±10 percentage-point equivalence margin while the sixth is borderline under that margin, and no model×memory interaction is detected. Two other headline quantities depend on analysis scope: a mitigation instruction's apparent harm to one model reverses sign once exploratory dimensions are excluded, and the share of unsafe turns that flip to safe when memory is masked falls from 42–76% to 26–43% when restricted to the confirmatory safety dimensions. Longitudinal safety benchmarks should report rater-, aggregation-, and attribution-sensitivity alongside headline numbers; we release SafeDrift to make this practical.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.