Exact-Verifier Interventions Identify Future-Failure Risk for Language-Model Control
Abstract
Exact verifiers create a distinctive control problem for language-model systems: a controller must commit, branch, abstain, or escalate before a symbolic grader, test suite, proof kernel, or state checker reveals the terminal outcome. We formalize control-object identification, a matched intervention that fixes the decoder policy, verifier, branch pool, feature interface, calibration split, optimization budget, and decision interface while varying the supervision target and decision functional. The resulting estimand is deployment-policy future-failure risk, the conditional probability that a branch-local action will eventually be rejected under the continuation policy. We show that this scalar yields the Bayes rule for local commit-versus-escalate decisions with a fixed fallback contract, while branch-set utility requires explicitly conditional joint failure risks. The paper supplies a read-only frozen-feature controller interface and a complete experimental record specification for testing competing targets without confounding them with data, architecture, calibration, or cost changes. This formulation turns late exact-verifier failures into a precise target-identification problem and establishes the controls required for credible comparisons of language-model control objectives.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.