AgentClutch: Verifier-Gated Computation Allocation for Stateful Agents
Abstract
An agent cascade can contain a test-passing fallback yet return an earlier failing output. We study this failure with AgentClutch, a fixed-order cascade that restores a registered execution state before fallback. Our analysis separates the success available among retained outputs, the success delivered by the stopping rule, and the executor cost of acquiring it. On 134 Python repair problems, replacing syntax stopping with public-example checks, while holding candidate and reference outputs fixed, raises delivered test success from 78.4% to 90.3%. The stronger gate recovers 16 blocked test-passing fallbacks at 29.1% higher executor-token use. In CooperBench and ScienceWorld, a middle executor adds unique operational acceptances but costs more than the downstream reference work its successful stops avoid. The full cascade nevertheless remains cheaper than standalone reference execution in the CooperBench census under a fixed token tariff. These findings show why complementary outputs, reliable selection, and economical acquisition must be evaluated separately. The studies concern restorable execution units rather than uninterrupted task completion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.