Learning from Local Repairs: Edit-Certified Advantage Redistribution for Rubric-Based Reinforcement Learning
Abstract
Rubric-based reinforcement learning provides richer feedback than binary rewards, but still broadcasts one response-level advantage across all generated tokens. On-policy self-distillation (OPSD) instead derives token-level guidance from distribution shifts induced by privileged context. However, privileged context can also alter wording and style, making distribution shifts an ambiguous basis for token-level credit assignment. To address this ambiguity, we introduce Edit-Certified Advantage Redistribution, which uses successful local revisions to obtain response-specific evidence for token-level credit. The failed criteria guide a reviser toward a local repair that preserves unrelated content. A revision is certified when it remains local, yields a correct answer, improves the weighted rubric score, and preserves previously satisfied criteria. Aligned edits and their neighborhoods define edit-certified repair support on the original response. ECAR uses this support to redistribute the original advantage while preserving its response-level mean and each token advantage's sign. Responses without such support retain the standard sequence-level RL update. ECAR achieves higher average performance than reinforcement learning and self-distillation baselines across five mathematical reasoning benchmarks. Ablations further establish the importance of both the revision certificate and edit-based localization. These results highlight the value of certified local repairs for token-level credit assignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.