From Self-Correction to Self-Evolution: Agentic Test-Time Learning for Mathematical Reasoning
Abstract
Recent advances in self-correction (SC) and recursive self-improvement (RSI) offer new ways to improve large language models' reasoning. A natural approach is to accumulate correction experience as reusable skills to guide subsequent problem solving. To this end, we propose SC2SE (Self-Correction to Self-Evolution), an agentic test-time learning framework for mathematical reasoning. With model parameters frozen and without requiring ground-truth answers as learning feedback, SC2SE learns from its own verification and revision experience through four collaborating agents for problem solving, stepwise correction, skill evolution, and role calibration. These agents distill correction experience into skills; role calibration checks each skill's suitability for the solver role before and after skill aggregation, turning diagnoses of existing solutions into actionable guidance for subsequent problem solving. Updated skills guide the next round of problem solving, and new answers undergo further verification and correction, forming a test-time self-evolution loop: correction → skill accumulation → skill-assisted answering → further correction. We conduct same-pool evaluation on four mainstream large language models and five public mathematics datasets. The results show that skill-assisted initial answers achieve higher mean accuracy (ACC) than the compared RSI baselines; with further correction, SC2SE improves mean ACC by 14.3 percentage points over the highest mean ACC among the compared baselines. Together with component comparisons and process case studies, these results support the complementary roles of self-correction and skill accumulation: correction provides experience for skill learning, while accumulated skills support subsequent problem solving and revision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.