acceptodds
Under review as a conference paper at ICLR 2027

RLSR: Reinforcement Learning with Self-Reflection for Large Language Models Self-Improvement

Abstract

Reinforcement learning (RL) can distill capacities that emerge with test-time self-reflection and incentivize large language models (LLMs) to self-improve. However, existing approaches fail to co-evolve self-reflection with direct policy optimization and instead condition the privileged self-teacher on unverified hindsight, causing self-improvement to plateau or even reverse. In this paper, we propose RLSR, a framework that optimizes and distills self-reflecting LLMs to achieve persistent self-improvement with RL. To generate endogenous learning signals, RLSR samples rollouts via grouped multi-turn self-reflection and retains only instances that achieve group-level aggregate success, grounding the privileged self-teacher in verified rather than assumed hindsight. RLSR also optimizes the policy via a dual-track co-evolutionary mechanism: a direct track internalizes breakthroughs incurred by self-reflection into the direct policy via sign-adaptive counterfactual advantage reshaping, and a reflect track assigns turn-level advantages via a log-odds belief calibration that tracks pivotal corrective transitions, actively reinforcing error correction to safeguard the self-reflection engine against degradation. These two tracks are optimized under a unified policy-gradient objective, establishing a sample-efficient self-reinforcing co-evolutionary cycle that remains stable throughout training. Across eight mathematics and code synthesis benchmarks, RLSR outperforms or matches RLVR and state-of-the-art self-distillation baselines in single-turn pass rate with evolved test-time reflection capacities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.