Identifiable Evaluation of Learned Systems: From Behavioral Collisions to Measurement Design
Abstract
Evaluation scores can be precise even when the protocol does not identify the behavioral property they are used to measure. We formulate protocol identifiability relative to a declared finite behavior class, an estimand, and a candidate measurement family. A protocol fails when two behaviors with different estimand values induce the same observable signature; such cross-estimand collisions are executable certificates of under-identification, and the collision structure itself defines a measurement-design problem: synthesize a minimum support that separates every declared alternative, or repair a failed protocol by minimum augmentation. The result is a constructive audit-to-repair workflow—expose the witness, add the cheapest distinguishing measurements, re-audit—assembled from classical hitting-set and concept-identification machinery. In a controlled language-model testbed with solver-verified labels, exhaustive audits show that base-only measurements collapse seven frozen policies, whereas a two-cell synthesized support separates them; on a fifty-policy stress class, certified synthesis identifies where 50,000 same-budget random supports all fail. Applied to classes derived from the models’ own predictions, a frozen 15-cell synthesized support identifies all 86 derivation signatures, while held-out re-audit exposes one residual collision and its one-cell repair; same-budget entropy and random supports leave substantially more unresolved alternatives. A confirmatory rerun with surface-matched target and sham edits then measures a positive base-conditioned target–sham asymmetry across four model families after controlling the prespecified edit-surface features. The framework turns a validity concern into an auditable procedure for measurement design; all empirical results are controlled-testbed consequences, not external validation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.