acceptodds
Under review as a conference paper at ICLR 2027

Faithfulness Above One: How a Circuit Benchmark Rewards Leaving Out Suppression

Abstract

MIB, the Mechanistic Interpretability Benchmark, ranks circuit-discovery methods by two areas of a faithfulness curve. On its grid, the higher-is-better CPR is a constant plus the overshoot of the curve above the full model minus its undershoot below it, and the lower-is-better CMD is their sum, so every published pair can be checked, and CPR rewards circuits for exceeding the model they are meant to explain. MIB computes CPR on circuits of the highest-scoring edges. For additive edge effects and exact scores this rule maximizes CPR, and its most faithful circuits contain no suppressive component. Among MIB's own entries, the leading CPR in 9 of its 11 model and task cells can only come from a curve above the model, and on GPT-2 small the CPR spread of five public submissions is overshoot. Removing the logit edges of the two negative name mover heads, part of the known IOI circuit, from the MIB example's magnitude-selected circuits makes them less complete and raises their CPR from 0.97 to 1.58, above the 1.24 of the circuits whose CPR MIB reports. Rescored under the ABC counterfactual, EAP, the MIB example's method, gives value-selected circuits that exceed the model at the sizes that carry most of the weight when evaluated under the answer flip that MIB runs and fall below it under ABC. The largest overshoot, MAttr's, is not explained by the two edges and arises in patched states far outside the model's range. We recommend computing both scores on one declared circuit family, with a per-example distance from the model as the headline behavioral score.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.