Coding Agents as Task Compilers: From Specifications to State-Centric Visual Control Specialists
Abstract
In visual control, a coding agent can build the controller rather than be the controller. Using a general-purpose coding agent as the online controller requires repeated model calls during each episode, incurring latency and inference cost. We instead concentrate coding-agent computation in development and formulate task compilation: compiling a task or task family into a reusable specialist that thereafter executes independently through a fixed runtime interface. We instantiate this idea with State-Centric Task Compilation (SCTC), which organizes the specialist around a task-state and memory module , a visual encoder , and a control policy . The state interface separates the information estimated from images, retained across time, and used for control, providing a compatibility boundary for goal-conditioned reuse and localized recompilation. Across the evaluated manipulation tasks and games, coding agents construct effective frozen specialists. In Craftax, a specialist compiled by Astra achieves similar mean reward to Astra playing the game directly, with about shorter episode runtime. In object pushing, updating only the visual encoder restores success under visual perturbations, while adding a generalization requirement to the specification improves performance on unseen shapes. In Minecraft, frozen specialists use screenshots and keyboard/mouse actions to craft a stone pickaxe within five minutes in unseen worlds. These results show that coding agents can turn task specifications into reusable visual control specialists with low execution overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.