acceptodds
Under review as a conference paper at ICLR 2027

Conditionally Approved: Proof-Bound Branchable Universes for Long-Horizon AI Agents Under Changing Evidence, Authority, and Time

Abstract

We study a methodological problem that static benchmarks cannot represent: an agent acts while its world changes, yet the evaluator must prove what changed, when the agent could observe it, which authority controlled it, and whether later work used current state. We introduce an instrument for constructing and evaluating proof-bound, evolving, branchable agent Universes, instantiated in a fully synthetic mortgage testbed. Five long-horizon tasks isolate document truth under pushback, epistemic residue after evidence retraction, conflicting legitimate authority, temporal portfolio control, and a handoff measured through successor usability. Every eligible episode is bound to a frozen task, model identity, runtime, action ledger, material-world roots, agent-visible projections, event and restore lineage, evaluator edition, and infrastructure disposition. Official reward is a strict binary conjunction; criterion vectors are diagnostic evidence rather than hidden partial credit. A post-observation cost-stop amendment capped an original pass-at-five plan at 55 physical attempts across four identity-pinned providers. Thirty-seven attempts were gradable and 18 were infrastructure exclusions. None of the 37 satisfied the perfect-rubric conjunction. Only six of 20 provider-by-task cells contained an eligible same-runtime pair, so the verifier refuses balanced provider comparisons, pass-at-five, and tail reliability claims. The ledger nevertheless exposes a repeatable two-action branch-interface short-circuit, 13 low-engagement traces, three runtime cohorts, and 889 red criterion surfaces packaged for a separately scoped blinded-attribution study. A post-hoc, advisory stress test then asked four independently pinned model judges to disposition all 889 red surfaces from the same frozen evidence. Only 232 assignments (26.1%) were unanimous; descriptive Fleiss' κ was 0.151, pairwise exact-label agreement ranged from 35.4% to 69.5%, and pairwise evidence-anchor Jaccard ranged from 0.089 to 0.292. No judge vote changed a frozen score. The completed instrument study shows that dynamic episodes can be frozen, evolved, restored, branched, and audited without collapsing infrastructure missingness or diagnostic evidence into model failure. Its claim is methodological; it is not a claim of frontier difficulty, mortgage validity, causal effects, provider ranking, or human agreement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.