acceptodds
Under review as a conference paper at ICLR 2027

ObserverBench: Testing Mechanistic Estimates For Intervention and Control

Abstract

Interpretability methods increasingly guide concrete actions, such as editing activations, ablating components, or allocating scarce human review. Yet statistical prediction accuracy alone does not guarantee effective downstream decisions. We present ObserverBench, a framework that evaluates an internal estimator—an *observer*— through the intervention, controller, or policy it informs. Each task fixes the available measurements, candidate actions, decision policy, and loss, reporting prediction error and action loss separately. Control analysis shows that, once an observer agrees at the starting state, it need only match the true response along directions the actuator can reach. Across empirical benchmarks, statistical quality repeatedly decouples from decision outcomes. In model editing on GPT-2 and Qwen2.5-7B, modeling component interactions improves effect predictions, yet lower mean error fails to select better interventions; predicting decision loss directly does. In safety triage, perfect classification can misallocate a fixed budget by ignoring consequence severity. On code-backdoor detection (APPS) across Qwen (2.5-7B, 3.5-9B) and Gemma-2-9B-it, residual observers yield the lowest average loss, but the optimal measurement context varies by model, and AUROC rankings can invert operational performance. ObserverBench establishes observer adequacy as a declared, task-relative, and reproducible evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.