acceptodds
Under review as a conference paper at ICLR 2027

Decision-Sufficient Epistemic Evaluation: When to Redesign and When to Acquire

Abstract

AI evaluation often optimizes additional measurement before asking whether measurement is a valid repair: when a representation erases an action-relevant distinction, more data through that representation cannot recover it. We formalize this prior repair question by distinguishing semantic erasure from statistical insufficiency through two repair gates. Semantic aliasing induces unavoidable intervention regret, while preserved distinctions may remain statistically unrecoverable. Locally, decision recoverability requires and, when estimable, is governed by . After preservation, classical directional design targets acquisition along the downstream decision direction. Across 2,400 held-out controlled interfaces with hidden failure mechanism and encoder matrix, a blind gate reaches 92.6% diagnosis accuracy and lowers post-routing regret to 0.1175 versus 0.2085 for Always Acquire and 0.1710 for Always Redesign. We further train 1,800 representations without assigning semantic/statistical failure labels. Using only independent calibration separability, the gate attains regret 0.1442, below Always Acquire (0.1480) and Always Redesign (0.1648). Post-hoc decision-subspace overlap tracks the repairability crossover. Results support a contract-specific criterion for deciding whether additional measurement is a valid repair under controlled populations, rather than a universal natural-failure classifier.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.