Beyond Spurious Reflection: Strategic Self-Correction as State Replacement in LLMs
Abstract
Large reasoning models excel by scaling inference-time Chain-of-Thought, yet their performance is brittle: an early mistake corrupts all subsequent reasoning. Reinforcement-learning-trained reasoners often produce reflective phrases like "wait, let me reconsider," but growing evidence shows these reflections are largely spurious–they neither alter reasoning paths nor final answers. Existing self-reflection methods keep the faulty chain in context, forcing the model to overcome its own errors. We propose REFLECT+, an end-to-end framework that directly optimizes reflection-triggered self-correction. Building on the Clean-Slate Reflection paradigm–where the model emits a structured <reflection> block summarizing verified sub-conclusions, diagnosed mistakes, and corrected directions, then continues from the problem and reflection alone–REFLECT+ adopts a two-stage training scheme combining supervised cold-start and role-aware reinforcement learning. This enables strategic decisions on when to reflect, what to retain, and how to resume. Experiments on Qwen2.5-Math-1.5B / Qwen3-0.6B-Base / Qwen3-4B-Base across MATH500, AMC23, GSM8K, LogiQA, GPQA, MMLU-Pro, BBH, CLUTRR and DROP evaluate both self-correction quality and reasoning accuracy, and outperforms GRPO on accuracy with selective reflection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.