From Traces to Guarded Programs: Evidence-Gated Compilation of Recurrent Agent Workflows
Abstract
Tool-using agents often invoke a model at every step, even along recurrent read-only paths. Replacing repeated reasoning with deterministic execution could reduce inference cost, but recurrence alone does not make substitution safe: tool arguments may depend on prior observations, apparently read-only operations may hide effects, and replaying the same calls may still alter the final answer. Recurrence identifies an optimization opportunity, not permission to remove reasoning. We introduce guarded agentic compaction (GAC), an evidence-gated approach that compiles recurrent execution paths into deterministic programs only when substitution is justified. GAC compiles initial read-only prefixes only: it grounds tool arguments in observable state, treats writes, approvals, handoffs, and unknown effects as hard barriers, protects compiled execution with input, output, and environment guards, and admits specialization only when finite-sample evidence satisfies configured risk and confidence requirements. Otherwise, it retains the largest justified prefix when possible or falls back to the unchanged agent. Admission certificates are exact finite-sample bounds under i.i.d. calibration groups, compiler-wide where one candidate reached calibration and per-candidate where two did (two of three primary families, whose corrected bound is 0.057). Across three live GitHub workflow families and 90 unseen test cases, GAC satisfies 90/90 exact task contracts and the unchanged agent 89/90, a difference of one stochastic baseline miss, while reducing model requests by 66.6%, tokens by 63.1%, latency by 64.2%, and estimated cost by 58.7%. A time-forward evaluation across five frozen public repositories specializes four and refuses one. On four public trace benchmarks, NESTFUL, API-Bank, and BFCL v4 contain recurrent, replayable traces but insufficient qualifying evidence to satisfy GAC's default 0.05 risk bound at the required confidence, whereas AppWorld provides sufficient evidence for admission. On a second model family and a second provider, the unchanged pipeline re-derives the same artifacts from each model's own traces on two of the three families and refuses the third. GAC reframes agent optimization as evidence-gated specialization: compile only what is justified, dispatch only when runtime conditions remain valid, and preserve model reasoning everywhere else.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.