Authenticated Late-Bound Evaluation of Persistent Agent Artifacts
Abstract
Persistent Agent artifacts can parse after an SDK upgrade yet change a later consumer's result. We study authenticated late-bound evaluation: an origin authenticates a complete historical view; a receiver binds a later consumer, context and interpreter, then compares historical and current results. Omitting consumer-observable distinctions precludes exact evaluation. We implement signed snapshots, proof-backed consumer execution (PBCE), and authenticated keyed events. Closed-IR simulation and authenticated openings yield conditional, IR-relative soundness; materialization and current execution remain trusted. Unsupported semantics abstain. Source checkers cover five AutoGen schemas and a corrected LangGraph reducer under explicit projection/conversion premises. An unfiltered AutoGen inventory certifies only 5/22 workloads. An empirical AppWorld confirmation covers 18/36 tasks with a frozen replay profile; retrospective repairs reach 36/36, but the Agent solves only 15/36. These empirical extensions remain outside the closed-IR theorem and do not establish independent transfer. Full snapshots favor small dense objects, and keyed events can be cheaper than PBCE. On 35 restored development checkpoints, all 2100 local receiver queries agree. Including selective sender work but granting prepacked full archives, task-median first-handoff time is 0.157 a compact full snapshot and 0.114 a root-authenticated lazy full ZIP; warm-query time is nearly equal. These are local handoff costs, not whole-Agent speedup or superiority over all lazy stores. Our contribution is a scoped evaluation lifecycle and measured coverage–cost frontier, not arbitrary-code certification, generic-store superiority or ecosystem-wide safety.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.