acceptodds
Under review as a conference paper at ICLR 2027

Do Not Go Gentle into Repetition: Escaping Overthinking with Repetition-Aware Reinforcement Learning

Abstract

Recent reasoning models achieve strong performance through reinforcement learning, but often generate excessively long traces containing redundant verification and repetitive loops. Existing efficiency-oriented methods mainly rely on length rewards or position-agnostic truncation, which may interrupt productive reasoning and can even assign negative advantages to correct trajectories. We propose , petition-aware reinocement lerng for efficient reasoning. REFRAIN uses repetition patterns to identify unproductive reasoning and construct shorter training rollouts. It further introduces correctness-preserving optimization to avoid distorting the learning signal of successful trajectories. Experiments across mathematical reasoning benchmarks and model scales show that REFRAIN improves both training and inference efficiency while maintaining or improving reasoning accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.