Content-Matched Is Not Lexically Matched: Text and Untrained-Network Baselines for Evaluation-Awareness Probing
Abstract
Linear probes are increasingly read as evidence that language models represent whether they are being evaluated. Content-matched designs, which hold a task fixed and switch one cue at a time, are meant to make that evidence strong. The assumption is that a probe separating the two versions has detected the model recognising the cue, rather than the cue itself. We test that assumption under a pre-registered protocol on 17 open models. Matching content is not matching words: a bag-of-words classifier given the probe's labels and splits separates all eight cues of EvalAwareBench (median AUROC 0.89–1.00); replacing "Marcus Chen" with "John Smith" is the manipulation, and a word counter sees it. The standard per-cue probe test does not need learned weights either: 13 untrained networks pass it on every cue and also separate a mixed-source evaluation/deployment set. Probe success on such benchmarks therefore cannot, by itself, show that a model represents being evaluated. Two results survive both baselines. Activations add a small but robust increment beyond the words, larger in trained than in untrained networks. And when both conditions use unseen words, probes on trained networks, like pretrained sentence encoders, still separate real from placeholder entities, while word models and untrained networks do not: they read a learned property of the cue, not evaluation. We release lexcheck and recommend a same-split text baseline and an untrained-network baseline for every probe claim.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.