acceptodds
Under review as a conference paper at ICLR 2027

Learning from Error Structure: Selective Supervision for Reasoning Post-Training

Abstract

Reinforcement learning with verifiable rewards (RLVR) and supervised finetuning provide complementary signals for reasoning post-training. As recent approaches increasingly combine them, a key question is where additional supervised correction is most valuable. Existing approaches often rely on rollout accuracy, using failure frequency as a proxy for supervision need. Our empirical analysis shows that this reliance on rollout accuracy can mask important differences in failure structure. Even at the same rollout accuracy, errors may either repeatedly collapse onto the same semantic mode or span diverse reasoning strategies. Motivated by this observation, we introduce LESS (Learning from Error Structure: Selective Supervision), which allocates auxiliary supervision according to recurrent failure structure. LESS groups incorrect rollouts by their solution method and final answer, measures error concentration with semantic entropy, and selectively applies verified reference demonstrations to recurrent-error questions while retaining standard GRPO for all rollouts. Experiments with Qwen2.5-Math-7B show that LESS significantly improves performance relative to existing methods across multiple mathematical, scientific, and broader reasoning benchmarks. These results suggest that effective supervision depends not only on how often a model fails, but also on how it fails.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.