acceptodds
Under review as a conference paper at ICLR 2027

Can LLMs Learn to Reason Robustly under Noisy Labels?

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) effectively trains reasoning models that rely on abundant perfect supervision, but its vulnerability to unavoidable noisy labels remains underexplored. Through empirical validation, we find that the vanilla RLVR method demonstrates degraded performance for two reasons: (i) data waste: when the LLM itself cannot roll out the noisy label, many data points are overlooked in training; (ii) policy misguidance: if the wrong label is rolled out, the policy will be guided toward the wrong direction, hurting model training. Fortunately, by inspecting the training dynamics, we identify an Early Correctness Coherence phenomenon: in the early stage of training, the probability of correct answers rises similarly for clean and noisy samples. Inspired by this, we develop a novel Online Label Refinement (OLR) algorithm that performs self-consistent label correction when a sample's predicted labels are historically stable and increasingly confident. Evaluated on six in-distribution (ID) math benchmarks and three out-of-distribution (OOD) tasks under 0.1–0.9 noise ratios, OLR achieves average gains of 3.6% (ID) / 3.3% (OOD) under inactive noise and 3.9% (ID) / 4.6% (OOD) under active noise.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.