Using Experience Twice: Dynamic Process Rubrics for Privileged Token-Level Credit Assignment
Abstract
Agent trajectories often contain both critical successful decisions and failed attempts; therefore, optimization with sparse trajectory-level rewards struggles to indicate which intermediate decisions to reinforce or suppress. We observe that meaningful distinctions among on-policy trajectories can be distilled into natural-language process rubrics and used as process supervision signals. Based on this insight, we propose Online Rubric Self-Distillation(ORSD), which leverages self-derived process rubrics for dense token-level supervision, thereby augmenting outcome-centric reinforcement learning. Specifically, at each update, the current policy independently generates multiple rubric bundles from the same group of trajectories, extracting task-specific candidate process criteria from their cross-trajectory behavioral differences. A frozen verifier then evaluates all rubric criteria against each trajectory in the group, and selects one candidate bundle based on bundle-level alignment between rubric satisfaction and task success. ORSD treats the selected rubric bundle as privileged information during training, then compares the log-likelihoods of the original tokens with and without the rubric. The resulting likelihood gap provides token-specific process credit, which is added to the original trajectory-level GRPO advantage through additive advantage shaping. ORSD thus bridges trajectory-level outcome supervision and fine-grained process supervision, enabling token-level credit assignment using process signals. Extensive experiments on ALFWorld and WebShop show that ORSD accelerates policy improvement and boosts final task performance for both Qwen3-1.7B and Qwen3-4B under the same environment-interaction budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.