Proof-Valid, Yet Unsafe to Commit: Fail-Closed Evidence Gating for High-Impact Agents
Abstract
Consider a remediation agent holding a signed exception whose inclusion proof verifies under checkpoint \(r_0\). Checkpoint \(r_1\) can revoke that exception while the served subset still verifies under \(r_0\); suppressing a required safety constraint creates the same proof-valid, action-unsafe commit. This trace motivates action admissibility and ProofLatch, which preserves a typed evidence capability from retrieval to actuation and releases an action-bound token only when fragment integrity, client-observed freshness, policy compatibility, and policy-bounded completeness hold at the commit boundary. On \(50,000\) matched attacked remediation/Qwen2.5-32B opportunities, proof-verifying fail-open consumption committed \(3,400\) compromised actions (6.8%), versus \(9,300\) (18.6%) without integrity; base and range-proof ProofLatch each committed \(0/50,000\) (one-sided 95% upper bound \(<0.0060%\)). The exhaustive base ledger records \(41,550\) immediate-safe, \(6,100\) delayed-then-safe, and \(2,350\) blocked outcomes, exposing the liveness cost rather than removing it from the denominator. Fail-open compromise remains 5.6–6.8% across two workflows and two model families while ProofLatch records none; separately, all \(37\) enumerated routes cross the gate, with no unauthorized side effect in \(118,400\) integration-fault injections or \(11\) held-out route/fault compositions. Base enforcement adds 8.7% attacked wall time, and range proofs trade 10.2% cost for 1,750 additional immediate-safe and 500 additional eventual-safe outcomes. The result identifies the commit boundary—not proof verification alone—as the control point that turns mutable retrieved evidence into authority.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.