acceptodds
Under review as a conference paper at ICLR 2027

Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning

Abstract

Reinforcement learning (RL) for large language model (LLM) post-training relies on scalar outcome rewards that convey only *whether* the model succeeded, not *why* it failed, forcing blind exploration that stalls in failure-dominated regimes where advantages collapse. Yet rich natural language (NL) feedback, from execution traces to textual critiques, often arises as a byproduct of the reward system itself but remains largely discarded. We propose GOLF, an RL framework that leverages **G**r**O**up-level **L**anguage **F**eedback to operationalize refinement as a learnable *meta-capability*. Rather than refining each failure in isolation, GOLF aggregates external critiques and intra-group attempts into a unified context, surfacing systemic error patterns that broaden solution diversity. The resulting refinements are adaptively injected as off-policy scaffolds in failure-dominated regimes, while generation and refinement are jointly optimized so that improved self-refinement yields stronger scaffolds that accelerate exploration. Across verifiable and non-verifiable tasks, GOLF achieves the best average score in every setting: it surpasses the strongest NL-feedback baseline, Critique-GRPO, by +11.86 points on Llama-3.1-8B-Instruct, and its gains hold consistently as Qwen3 policies scale from 4B to 14B. It also matches the baseline's final performance in up to fewer steps, while sustaining broader exploration and cultivating test-time self-refinement that generation-only RL fails to develop.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.