Same Task Performance, Different Intervention Readouts: Cross-Run Variation under In-Context Rule Aliasing
Abstract
Causal edits can yield different readouts from models with nearly identical task behavior. We study 26.1M-parameter transformers on a synthetic assignment language where RECENCY and RARITY select the same answer on every training document. All 75 runs across 25 configurations reach accuracy at least 0.999, yet under an answer-preserving multiplicity edit, the sign fraction—the fraction of valid pairs with positive log-odds displacement—spans more than 0.3 across seeds in 13 cells, reaching 0.879; there a strong positive tail contrasts with mostly near-zero negatives. Replacing the probe batch moves this fraction by at most 0.034 in four extreme runs. These are differences in intervention sensitivity rather than answer policy: on generated documents where the rules disagree, rarity-answer rates are at most 0.003 in the main grid. Supervision that distinguishes the rules changes the readout. A RARITY-labeled arm supplies a one-sided trained control for positive displacement; in the RECENCY-labeled arm, observed dispersion is lower. The dispersion comparisons remain exploratory and are not adjusted for cell selection. Strong directional readouts occur before runs meet the retrieval diagnostics. Intervention readouts should be reported with behavioral tests, independent training runs, checkpoint, probe population and gating policy rather than treated as configuration-level mechanism labels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.