acceptodds
Under review as a conference paper at ICLR 2027

SETA: Amortizing Test-Time Compute with Explicit-Failure Memory for Tool-Using Agents

Abstract

Agents can fail a task and still judge themselves successful. Success-only memory discards this signal, and K-attempt self-consistency multiplies compute. SETA (Selective Explicit-failure Trajectory Amortization) turns externally verified failures into reusable memory. Offline, we collect four exploratory trajectories per task and flag false successes. We deterministically rewrite each: the error path stays, only the outcome changes, with no model call. At deployment, a single attempt retrieves this memory: a success skeleton, known-pitfall exemplars, and one explicit-failure memory. For example, a wrongly approved leave request becomes a note the agent reads when the same request returns. We evaluate 60 requesting-time-off tasks and 50 customer-request-routing tasks from AgentArch, assuming mature memory for recurring requests; cold-start and cross-task transfer are out of scope. Under this assumption, a single Think-mode attempt reaches 85.00% final-outcome success on time-off tasks. On a separate seed set, removing only the explicit-failure memory from otherwise identical retrieval lowers final-outcome success by 16.00 points. A corrected bf16 harness reproduces this effect at 15.67 points; a second stack shows 9.00 points, and retrieval alone gives no gain on either. In Think mode, SETA exceeds a workflow-memory baseline, an LLM-written lesson-memory baseline, and four-attempt majority voting without memory. SETA and the workflow-memory baseline pay the same strict-task-success cost, but only SETA converts it into a gain; under non-thinking (NonThink) deployment, both are indistinguishable from no memory. On the longer routing tasks, memory injected once has no effect. Restating explicit-failure memory each call yields a 19.20-point gain; restating retrieval-only memory yields none. 5

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.