Cogito, Ergo Ludo: Learning External Language States for Interactive Agents
Abstract
Reinforcement learning can produce effective behavior from interaction, but the knowledge it acquires is often stored implicitly in model parameters rather than as inspectable rules or strategies. We introduce the External Language-State Agent (ELSA), an LLM-agent framework for prompt-level unknown-rule interaction, where task names, hand-written rules, and demonstrations are not provided in the prompt. ELSA represents learned rules and strategies as an external language state initialized without task-specific content. During each episode, this state is fixed and conditions a language-based value function, a language-based world model, and action selection. Across episodes, post-episode reflection revises the external state, while GRPO trains the LLM to better use the evolving state from sparse outcome feedback. Across four symbolic, text-observable environments, ELSA improves over name-free prompting and static rule-given prompting controls while producing inspectable external-state artifacts. Experiments and controls show that these gains depend on iterative external-state refinement and learned use of that state, rather than static prompting or policy optimization alone. These results suggest a route toward interactive LLM agents whose task competence is accompanied by explicit, inspectable, and reusable language-level knowledge.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.