When All Rollouts Fail, Critique Provides Direction
Abstract
Group-relative reinforcement learning is widely used to improve reasoning in large language models. However, it cannot learn from a problem when all sampled answers are wrong, because their relative advantages collapse to zero. Critique-guided revision can continue exploration beyond these failed answers. Yet when all revisions are also wrong, the learning signal disappears again because outcome rewards cannot distinguish useful critiques from unhelpful ones. We propose Signpost, a training framework based on Critique as Direction. A critique can be useful even if it does not immediately produce a correct answer, as long as it points the model toward one. Signpost trains the same language model to generate both answers and critiques. Its directional reward measures how much a critique increases the likelihood of a target solution under a frozen reference model. This provides a learning signal when all revisions fail, without rewarding incorrect answers. Once a critique helps discover a verified solution, Signpost trains the model to produce that solution directly from the original problem using a bounded-likelihood objective. This transfers critique-guided discoveries into question-only solving and closes the gap between training and inference. Experiments on mathematical reasoning show that Signpost outperforms GRPO, DAPO, Scaf-GRPO, and Critique-GRPO. On Qwen3-8B, it improves pass@8 over GRPO by 6.6 to 10.0 percentage points across AIME 2024–2026. Critiques are used only during training, while inference requires no critiques, revisions, or additional computation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.