acceptodds
Under review as a conference paper at ICLR 2027

Beyond Spurious Reflection: Strategic Self-Correction as State Replacement in LLMs

Abstract

Large reasoning models excel by scaling inference-time Chain-of-Thought, yet their performance is brittle: an early mistake corrupts all subsequent reasoning. Reinforcement-learning-trained reasoners often produce reflective phrases like "wait, let me reconsider," but growing evidence shows these reflections are largely spurious–they neither alter reasoning paths nor final answers. Existing self-reflection methods keep the faulty chain in context, forcing the model to overcome its own errors. We propose REFLECT+, an end-to-end framework that directly optimizes reflection-triggered self-correction. Building on the Clean-Slate Reflection paradigm–where the model emits a structured <reflection> block summarizing verified sub-conclusions, diagnosed mistakes, and corrected directions, then continues from the problem and reflection alone–REFLECT+ adopts a two-stage training scheme combining supervised cold-start and role-aware reinforcement learning. This enables strategic decisions on when to reflect, what to retain, and how to resume. Experiments on Qwen2.5-Math-1.5B / Qwen3-0.6B-Base / Qwen3-4B-Base across MATH500, AMC23, GSM8K, LogiQA, GPQA, MMLU-Pro, BBH, CLUTRR and DROP evaluate both self-correction quality and reasoning accuracy, and outperforms GRPO on accuracy with selective reflection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.