acceptodds
Under review as a conference paper at ICLR 2027

Breaking the Solver Bottleneck: Training Problem Generators at the Learnable Frontier

Abstract

The limiting resource for training agents via reinforcement learning (RL) is increasingly frontier task supply: valid, solvable tasks just difficult enough to train the current model. As reasoning and agentic models improve, fixed task distributions saturate, while naive synthetic generation yields tasks that are trivial, impossible, or ill-posed. Training a task generator with RL to optimize validity and learnability can address this bottleneck, but direct optimization requires repeated solver rollouts per candidate. For software-engineering (SWE) tasks, a single rollout can take tens of minutes, and solver-in-the-loop generator training needs several such rollouts for every candidate task at every update. We introduce **PROPEL**, a solver-amortized framework for training task generators at the targeted solve rate. PROPEL trains a lightweight activation probe on a one-time labelled corpus of generated tasks and solver outcomes. The probe predicts target-solver pass rate from a frozen generator reference model and serves as a proxy for solve rate during generator optimization, reducing generator evaluation to a single forward pass. Across math, code, and software-engineering at multiple model scales, PROPEL shifts generation toward the targeted solve rate: for coding, tasks generated at the learnable frontier increase from %% for a *Qwen2.5-3B-Instruct* solver and from %% for a *Qwen2.5-7B-Instruct* solver. For SWE, PROPEL increases targeted solve rate generations from %% for *Qwen3.5-27B* on repositories not seen during training of probe and generator.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.