acceptodds
Under review as a conference paper at ICLR 2027

A Probe Direction Is a Property of Its Prompt

Abstract

The prompt used to build an activation probe for evaluation awareness decides the sign of the scaling result that the probe reports. Such a probe contrasts a model's activations on prompts framed as a test with its activations on the same prompts framed as ordinary use, and scores how well the resulting direction separates held-out items. We keep the task text byte-identical and cross a set of test framings with a set of deployment framings, on four series of open-weight models, each released at several sizes, up to the current Qwen3.5 generation. On every series some pairs of framings reproducibly make the score rise with model size and others make it fall, and this one design spans both published scaling results on this benchmark, one positive and one flat. The model being measured accounts for a point estimate of at most 13.1% of the variance in the score, while the way each model responds to each pair of framings contributes at least twice as much as the way it responds to each evaluation item. A reliable comparison between models therefore averages the score over many pairs of framings rather than adding items, and we measure how many are needed by checking that averages over disjoint sets order the models alike. A direction that carries no information about evaluation also reproduces much of each published score, because the benchmark's two classes differ in surface form, so every score should be read against that baseline. Code: https://anonymous.4open.science/r/probe-direction-384B/

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.