acceptodds
Under review as a conference paper at ICLR 2027

AGAR: A REINFORCEMENT LEARNING SUBSTRATE FOR LLM PROGRAM EVOLUTION

Abstract

Given a task and an evaluator, a language model can rewrite a candidate program while a search loop decides which rewrites survive — a practical route to algorithm discovery. But that loop is governed by five constants set by hand — which parent to select, how hard to mutate, how to keep diversity, what to remember,and a scalar score that never says which part of the program earned it — and reinforcement learning already has an estimator for each. The obstacle is that program evolution is not usually written down as a decision process. We formalize it as a Markov decision process whose action is the modular prefix the model is conditioned on, not the program it emits, so that credit assignment, value estimation,adaptive exploration and experience memory each attach to a distinct component.AGAR (Algorithm Generation As RL) provides the substrate this licenses: any estimator can be replaced or switched off without touching the controller, making the transfer auditable one mechanism at a time, with no gradient training of the backend model. Across 19 tasks, two backends and three seeds under one harness,AGAR improves on the stronger of two published baselines on most tasks, with the gains concentrated on the competitive-programming family. The formalization also yields a checkable reading of prior work: these systems are implicitly zero-discount, not by choice but because fitness is exogenous to an individual rather than a return over successors, leaving a discount factor nothing to act on.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.