Looking Like a Test: What Format Does to Evaluation Awareness in Language Models
Abstract
Language models often recognize when they are being evaluated, which can make benchmarks unreliable. Yet it remains unclear why a model identifies a conversation as an evaluation: whether format plays a part, and whether changing it changes behavior. To study this, we use judge models from eight families to compare pairs of conversations from evaluations and from deployment and choose the one that seems more like an evaluation. A pairwise preference model turns these choices into a single scale of evaluation-likeness, on which the judges largely agree. We then find what drives these choices by marking each conversation for hand-defined format cues and ranking the cues by how far they move the score. A question that comes with answer options looks like a test to the models, while a conversation with several back-and-forth user turns looks like real use. Because this ranking shows only an association, we rewrite hazardous-knowledge questions into versions that change format one rewriting step at a time. As a result, the question wrapped in exam framing scores higher than any evaluation conversation, while the same question inside a realistic work conversation scores lower than nearly every deployment conversation. Finally, we instruct models in the system prompt to underperform, then give them the hazardous and benign questions rewritten with each format cue. With refusals set aside, the drop in accuracy tracks the evaluation-likeness score of each version in most settings, including models that never served as judges. Together, these results suggest that format shapes how evaluation-like a conversation appears to models, and that presenting evaluation questions in realistic conversations may bring model behavior on benchmarks closer to behavior in deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.