acceptodds
Under review as a conference paper at ICLR 2027

Think in Language, Act in Code: Searching Programmatic Policies via Natural Language

Abstract

Programmatic policies offer an appealing representation for reinforcement learning because they can translate the general knowledge encoded in foundation models into functions that issue low-level actions in the environment. However, synthesizing effective programs through interaction with an environment is challenging, as the space of possible programs is vast and small syntactic changes can lead to large and often unpredictable changes in behavior. We tackle this challenge by searching in an abstraction of the program space through LLM-proposed natural language programs, using feedback from environment interactions to guide the search. Each natural language program specifies a candidate strategy and is translated into an executable policy through iterative Python code synthesis. This creates a two-level search process in which natural language guides the exploration of strategies while program synthesis grounds these strategies into executable policies. Experiments on a large set of MiniHack tasks show that searching through natural language programs improves both the efficiency of programmatic policy search and the quality of the policies discovered.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.