POLO: Persistent LLM-Objects with Self-Authored State
Abstract
The dominant paradigm for building software with LLMs is agent-assisted code generation. We ask not whether LLMs can produce software, but whether they can be it. We propose Programmable Online LLM-objects (POLO): software as a set of persistent LLM-objects that own domain responsibilities, communicate through natural language messages, and are created and modified in natural language while running. The paradigm's core proposition is a new memory discipline: an LLM-object holds structured, private, and self-authored state with the semantics of program variables. The object derives its state schema from its natural language definition and can revise it when that definition changes. Each message handling rewrites the values its rules need, keeps dated entries only where a rule counts over a window, and discards the processing trace. Handling depends only on the object's definition, its state, and the incoming message, however long the system has been running. We build FlowStateBench, a benchmark of stateful, long-horizon workflows spanning weeks to months of simulated time, where correct handling of each event depends on accumulated data across the entire history of the workflow, and on rule changes arriving mid-stream. On FlowStateBench, POLO outperforms agent frameworks representing four memory strategies (LangGraph, Letta, CrewAI, and OpenClaw), and replacing its state with a transcript in the same runtime lowers state-dependent accuracy by 22.9 points. On Memora, an external long-term memory benchmark, a self-authored state schema performs comparably to one written by a developer, and both score above our single-agent baselines. The code, benchmark, and generator are released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.