acceptodds
Under review as a conference paper at ICLR 2027

Trajectories as Programs: Verified Replay and State Recovery for LLM Agents

Abstract

Large language model (LLM) agents are typically given multiple attempts at the same task. A failed attempt has often already completed part of the task. To reuse past attempts, existing methods turn them into text and ask the model to regenerate the actions from the initial state. The regenerated actions may be different from the original ones, so text cannot bring the task exactly back to a state the previous attempt reached, and the completed part cannot be kept. To address this problem, we propose Trajectories as Programs (TaP), a training-free framework that stores each attempt as an executable program, so a new attempt can start exactly where the completed part ends instead of regenerating everything from text. To decide where a new attempt should start, we introduce Verified State Localization: it replays the program of a failed attempt and keeps the states that have verifiably completed part of the task as recovery points. But recovery points are not equally useful, and attempts are limited. To spend the limited attempts on the recovery points most likely to solve the task, we introduce Probe-Guided Attempt Allocation: cheap probes first judge which recovery points can leave the old failed path and make new progress, and full attempts go to the most promising recovery points. Once a task is solved, TaP replays its verified program directly instead of calling the model, so the task stays solved and every new attempt goes to the tasks that are still unsolved. Across AppWorld and BFCL-V3 and across model sizes, Tap sets a new state of the art on both metrics: it solves more tasks per attempt on average, and solves more tasks at least once over all attempts, than all baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.