acceptodds
Under review as a conference paper at ICLR 2027

Agent Behavior as Code: Efficient and Robust LLM Agents with Programmatic Specifications

Abstract

AI agents based on foundation models (FMs) have demonstrated strong capabilities to perform complex open-ended tasks. However, they face some common challenges in practice: (a) agent behavior can deviate drastically even for semantically similar tasks, leading to catastrophically propagated errors; (b) high cost and latency due to FM calls, repeated in full whenever a task recurs with different inputs; (c) FMs' limited context and instruction following capability confine how well agents manage the ever-growing execution context and follow complex plans. We introduce **A**gent **B**ehavior as **C**ode **Agent** (ABCAgent), which uses a symbolic program (\eg, Python code with potential neural functions) to fully specify the agent's behavior at runtime, with a powerful FM agent editing that program for flexibility. Behavior is thus specified without premature variable binding, and its execution is deterministic. We evaluate ABCAgent on six agent benchmarks, two of which we construct to test how well a derived program generalizes to variants of the task it was written for. ABCAgent matches a model-matched neural agent on GAIA and augmented GAIA, and surpasses it where robustness and long control flows matter: 98.3% against 97.3% on GSM-Symbolic (), 71.9% against 47.4% on the telecom domain of -bench (), and more records written correctly at every loop length on our control-flow-augmented WorkArena benchmark. For more parametric task families, ABCAgent is also significantly superior in efficiency. Without authoring a new program, ABCAgent solves 92.6% of GSM-Symbolic instances and 20.1% of augmented GAIA variants, which yields lower latency and lower cost on GSM-Symbolic, 19% lower cost on augmented GAIA, and lower agent latency on -telecom.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.