Belief Is Not Authority: Warranted State Transitions for Black-Box Language Models
Abstract
Guardrails check rules and state trackers infer latent needs, but neither decides when an uncertain belief may authorize a different action. We introduce RAVR-W (Retrieval-Augmented Verification and Repair with Warrants), a runtime-assurance layer for frozen black-box language models built on authority-flow semantics. Any discovered factor may enter an open shadow state, but it influences prediction, recommendation, or execution only through distinct, earlier-epoch receipts for each tier; retrieval supplies evidence, never authority. A typed transition warrant switches away from a known fallback only with fallback-relative effect evidence, binds every argument to its evidential lineage, and commits external effects only after a read-back witness. We prove authority non-interference (unauthorized factors cannot change the decision), tier non-escalation, fallback-relative switch validity, witness-closed commit, and preservation of learned preconditions, and we introduce blind causal transition induction, which grows the executable state schema from an empty graph without letting any episode authorize itself. Across 14 benchmarks in counseling, negotiation, and stateful tool use, against more than a dozen matched baselines, RAVR-W wins wherever evidence licenses a switch. It improves calibration on all six core counseling tasks with clustered intervals below zero (296,796 held-out instances; median Brier reduction 14.7%, up to 30.8%) and raises τ-bench state balanced accuracy from 0.585 to 0.726. On ToolSandbox it improves success in all four model/provider configurations (+0.0391 [0.0239, 0.0531]), on an independent 105-scenario set, and by +0.154 [0.050, 0.262] on 60 further scenarios; the complete live kernel with learned precondition closure improves a bare agent by +0.111 [0.052, 0.177] on 184 scenarios, and induced preconditions lift held-out tool-call success from 0.40 to 1.00. On Gaia2 it solves 5/6 workflows versus 0/6 for the default agent. Where evidence is insufficient, it returns the exact fallback (148/148 on MHDialog), turning refusal into an auditable outcome rather than a silent failure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.