acceptodds
Under review as a conference paper at ICLR 2027

ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement

Abstract

Reinforcement learning with verifiable rewards (RLVR) has advanced the reasoning capabilities of large language models (LLMs) to the point where frontier models match or exceed human experts on many tasks. Self-refinement, where a model reviews and revises its own solutions, can recover problems that its first attempt narrowly misses. However, prompt-only refinement remains brittle even for frontier models, and existing training-based methods rely on external guidance such as critique annotations or explicit signals indicating whether the initial answer is correct. In this work, we introduce ThinkTwice, a simple two-phase RLVR framework that jointly optimizes an LLM to solve reasoning problems and refine its own answers, without any external signals. Concretely, ThinkTwice replaces every even GRPO step with a refinement step, which optimizes the model to revise its own solutions from the preceding step under the same binary correctness reward, so training takes only 3% longer than GRPO. Across five mathematical reasoning benchmarks and two model families including Qwen3-4B and Olmo3-7B, ThinkTwice improves both reasoning and refinement performance over competitive online policy optimization baselines. Specifically, on Qwen3-4B, ThinkTwice outperforms GRPO on AIME by 5 percentage points before refinement and by 11.5 points after one self-refinement step. Analysis of the training dynamics of ThinkTwice reveals an implicit rectify-then-fortify curriculum: refinement predominantly corrects errors early in training and naturally shifts toward preserving already-correct solutions as the model improves, yielding a more rectified reward signal. Our work establishes joint training of reasoning and self-refinement as a principled and effective methodology for RLVR.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.