acceptodds
Under review as a conference paper at ICLR 2027

Which Environments Should LLMs Learn From? Predicting RL Utility Across Scale

Abstract

Reinforcement learning with verifiable rewards can improve language-model reasoning. Yet equal compute can help, do little, or degrade a policy across mathematical prompt distributions. We ask whether measurements from the starting policy can predict a distribution's held-out utility before observing its RL outcomes. We run 32 controlled Qwen3 trajectories from 1.7B to 14B parameters in a 22 design that varies group-signal availability and the admission-cost profile. We evaluate 196 policies on a fixed, disjoint 1,000-prompt OpenR1 evaluation panel. Environment choice changes held-out pass@1 by up to 11.4 percentage points at matched generated-action-token budget. Reward variation alone does not explain this spread. We measure admitted advantage mass per generated token, , and its effective coverage across prompts, . A compact - utility model is fit through 8B and frozen before inspecting 14B outcomes. At 14B, it achieves 2.29-point curve RMSE, 77.8% pairwise rank accuracy, and 11/12 cell-budget means inside its empirical envelope. A four-scale leave-one-scale-out audit yields 2.09-point RMSE. Fixed-label and crossover rules remain competitive for allocation, while a frozen uniform-mixture test identifies the current boundary of composition transfer. Together, these results provide evidence that pre-RL measurements combined with earlier-scale outcomes can help predict how environment choice changes held-out improvement across model scale and compute.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.