acceptodds
Under review as a conference paper at ICLR 2027

RECOMATH: Benchmarking and Improving Self-Correction in Mathematical Reasoning

Abstract

Self-reflection has emerged as a promising approach for improving mathematical reasoning in large language models (LLMs). A central capability of self-reflection is self-correction—the ability of models to identify and recover from errors in their own reasoning. However, existing mathematical reasoning benchmarks primarily evaluate whether models can solve problems from a clean starting point, providing limited insight into whether they can recover once an intermediate reasoning error has occurred. To address this gap, we introduce ReCoMath, a step-level benchmark for evaluating self-correction in mathematical reasoning. Each example pairs a problem with a reasoning prefix ending at the first erroneous step and asks the model to continue the reasoning without being told that the prefix contains an error. Our evaluation across several open-weight LLMs reveals that strong mathematical problem-solving performance does not necessarily translate into reliable self-correction. Using ReCoMath, we conduct a controlled study to isolate the effect of training-trajectory construction, showing that retaining intermediate reflection or repair steps consistently improves self-correction compared with trajectories that omit these processes. Moreover, correct-trace reflection yields more stable gains on easier instances whereas error-retaining repair is stronger on harder ones. Motivated by these findings, we propose Calibrated Reflection Training (CRT), which asymmetrically combines correct-trace reflection with a smaller proportion of error-retaining repair. Across three model backbones, CRT achieves a better overall balance between self-correction and mathematical problem-solving performance than either supervision signal alone. Code, data, and benchmark resources are available at https://anonymous.4open.science/r/RecoMath-4323.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.