acceptodds
Under review as a conference paper at ICLR 2027

SEER: Meta-Reinforcement Learning for Exploratory Search in Code LLM Post-Training

Abstract

Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving the reasoning performance of large language models. Yet its improvement is constrained by a simple bottleneck: RL can directly reinforce only reasoning trajectories that the policy has already discovered. Increasing rollout budgets can expose rare successes, but through increasingly expensive, undirected sampling; meanwhile, objective-level interventions reshape the policy distribution without specifying where or how the model should explore. Consequently, useful solution strategies may remain effectively inaccessible under practical rollout budgets. We address this limitation with SEER, a meta-reinforcement learning framework for Code LLM post-training targeting useful solutions rarely exposed under practical rollout budgets. SEER organizes attempts into meta-trajectories, where prior code, verifier feedback, and reflections guide search. Credit extends across attempts: an attempt is valued for immediate correctness and for enabling later success, so failed but productive attempts receive downstream credit. A meta-GRPO objective combines direct and cross-attempt signals. Yet search-conditioned discovery alone is insufficient: behaviors found under search context may remain inaccessible from the original prompt. SEER therefore transfers discoveries to the original-prompt policy via a distillation-augmented objective, enabling context-free inference. SEER improves mean Pass@1 by 3.58–5.40 percentage points over the baseline average; it reaches 33.8% coverage on previously unseen cases and gains +5.88 points in meta-search over StepCoder, supporting structured exploration and reintegration as an path for extending RLVR.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.