acceptodds
Under review as a conference paper at ICLR 2027

Replication Is Not Specificity: Negative Controls for Representation-Based Reward-Hacking Detection

Abstract

Representation-based detectors of reward hacking are validated mainly by accuracy and by replication across tasks, architectures or models, and only in part by controls like those we apply here. Replication shows that a statistic is stable, not that it tracks reward hacking rather than training progress, text style or ordinary change. We apply three controls to our own significant results: a clean-run placebo, a benign-onset control in which a genuine capability unlocks on the exploit's schedule, and a scorer swap that holds the text fixed and changes the model reading it. In nine toy environments, activation anisotropy shifted at hacking onset in 8 of 9, and as often in runs with no hacking. Against benign onsets it separates only one exploit type, a false self-report (AUROC 1.00 and 0.98 under an MLP and a GRU; 0.37–0.69 elsewhere), and there the raw observation stream separates the arms as well. In three LLM families, base models that never saw reward-hacking data reproduce 97.7% of the anisotropy gap between hacking and control completions. An onset detector fires on 8/8 hacking and 0/3 clean runs in two models, yet on re-collections in three models, clean fine-tunes and untuned base models reading the same text reproduce 97–110% of the rise. A hack-versus-clean linear probe on raw activations keeps a small model-state component under the scorer swap, but in two of three models a benign change of fine-tuning data leaves one too, and a bag-of-words reader of the same text comes within 0.021 AUROC of the probe. A planted model-state signal on matched inputs, which the scorer swap passes by construction, separates from benign runs at AUROC 0.57–0.63 at the smallest dose and 0.81–0.97 at the largest. We release the suite and all 357 runs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.