acceptodds
Under review as a conference paper at ICLR 2027

How Many Probes Identify What a Network Computes?

Abstract

Two networks can agree on every test example and still compute different functions. Benchmarks cannot separate them when the relevant disagreement is rare under the evaluation distribution. We show that any score formed by averaging a bounded loss differs by at most between two mechanisms that disagree on an fraction of that distribution, independent of benchmark sample size. Choosing the inputs is what helps. For a probe set with separation , identification cost scales inversely with , while a sufficient budget grows only logarithmically with the number of candidate mechanisms; per-query unreliability is controlled by a Bhattacharyya exponent . The predicted inverse-separation scaling is confirmed at slope () across a 25-fold range of . A greedy multicover identifies one of 256 mechanisms using 24 probes, while random draws from the same pool still misidentify 31% of models after 256 queries. On sixteen trained transformers, one designed probe reaches the 95% identification target that requires five random probes, and the random-probe curve is predicted without fitting. Across seven Pythia-410m checkpoints, candidates derived from the prompt and corpus statistics recover the checkpoint at which the dominant behavior changes. Across these settings, the value of probe design is governed by a separation quantity that can be computed before querying the model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.