acceptodds
Under review as a conference paper at ICLR 2027

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

Abstract

Frontier language models are trained to exhibit stable programming assistant personas, yet those personas may drift over long agentic coding sessions in production. After a long real-world coding session, a model that hedges preferences in a fresh chat (“I don't have preferences”) may instead express them directly (“Python—the feedback loop is instant…”). Existing persona-stability studies focus on scripted chat of tens of turns, leaving long production coding sessions involving thousands of tool-using agent steps, repeated compaction, and hours of interaction largely uncharacterized. We introduce ContextEcho, a benchmark and reusable harness with a 25-probe identity suite, a snapshot-then-probe protocol that evaluates conversation state without perturbing the main session, judge-based and judge-free metrics, and a corpus of 57 long agentic-coding sessions collected with consent and redacted for release across Claude Code and Codex CLI. We use one representative session for a 23-model cross-provider study and all 57 sessions for corpus-scale trajectory analysis across frontier and open-weight models. Results show that persona drift occurs across providers and emerges early in real agentic interaction, while in-session compaction neither reliably resets nor amplifies it. A single-shot anchor restores the original assistant persona, though placebo controls suggest a generic register-reset effect. On same-task continuations, drifted prefixes improve tool-call fidelity—which we attribute to retained task context rather than drift itself. In tool-free chat, however, the same prefixes break output-format contracts and inflate output length. These findings motivate ContextEcho as an open-source framework for auditing persona drift in long coding sessions without retraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.