acceptodds
Under review as a conference paper at ICLR 2027

ASSUMPTION-STRATIFIED AUDITING OF DETERMIN-ISTIC EVALUATORS

Abstract

Correct observed labels do not identify an evaluator’s behavior on valid counter-factuals, and a ffnite class of plausible faults is not exhaustive merely because it ffts those labels. We formulate small-budget auditing of deterministic evalu-ators as uncertainty over a fault’s witness set: the probes on which it disagrees with an independent semantic oracle. We separate three assumptions: a failure resembles a modeled survivor; some frontier probe reveals it; or multiple wit-nesses occur inside a visible contract mechanism. Our central result identiffes what optimal unstructured protection leaves unresolved. At budget two, it ffxes how often each probe is queried but not which probes are queried together, yield-ing an exact range of detection probabilities for the same multi-witness fault. The mechanism-supported assumption constrains part of that remaining freedom, and our allocator balancesthe registered assumptions. Acrosssix frozen IFBench eval-uators, the resulting cases include a free row-41 reffnement, costly reffnements for no whitespace and sentence:keyword, and frontier insufffciency under the modeled assumption for three tasks. Frozen hidden-fault validation is mixed

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.