AGAR: A REINFORCEMENT LEARNING SUBSTRATE FOR LLM PROGRAM EVOLUTION
Abstract
Given a task and an evaluator, a language model can rewrite a candidate program while a search loop decides which rewrites survive — a practical route to algorithm discovery. But that loop is governed by five constants set by hand — which parent to select, how hard to mutate, how to keep diversity, what to remember,and a scalar score that never says which part of the program earned it — and reinforcement learning already has an estimator for each. The obstacle is that program evolution is not usually written down as a decision process. We formalize it as a Markov decision process whose action is the modular prefix the model is conditioned on, not the program it emits, so that credit assignment, value estimation,adaptive exploration and experience memory each attach to a distinct component.AGAR (Algorithm Generation As RL) provides the substrate this licenses: any estimator can be replaced or switched off without touching the controller, making the transfer auditable one mechanism at a time, with no gradient training of the backend model. Across 19 tasks, two backends and three seeds under one harness,AGAR improves on the stronger of two published baselines on most tasks, with the gains concentrated on the competitive-programming family. The formalization also yields a checkable reading of prior work: these systems are implicitly zero-discount, not by choice but because fitness is exogenous to an individual rather than a return over successors, leaving a discount factor nothing to act on.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.