Perdura: Stateful Code Execution Harness for Long-Horizon Agents
Abstract
Long-horizon agents must generate code, invoke tools, and incorporate feedback across extended tasks. Many existing harnesses leave part of this execution continuity in the model-visible interaction history, requiring the model to recover prior outcomes, available computational state, and unfinished work as trajectories grow. We propose Perdura, a stateful execution harness in which the model chooses executable Python actions, while the runtime owns execution continuity. Across steps, the runtime retains intermediate computational state, exact records of prior operations, and unresolved work through a persistent namespace, an execution ledger, and an explicit work lifecycle. Each model call receives only a bounded context derived from runtime state and the current observation. Across three long-horizon benchmarks and four widely used agent harnesses under the same base model (GLM-5.3 Flash), Perdura attains task success comparable to the top-scoring harness on each benchmark while using 26–53% fewer input tokens than that harness, and no harness achieves both a higher score and a lower cost on any benchmark. Trajectory analyses show how computational reuse, registered waits, and recorded execution evidence support continuation across steps.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.