acceptodds
Under review as a conference paper at ICLR 2027

Using Experience Twice: Dynamic Process Rubrics for Privileged Token-Level Credit Assignment

Abstract

Agent trajectories often contain both critical successful decisions and failed attempts; therefore, optimization with sparse trajectory-level rewards struggles to indicate which intermediate decisions to reinforce or suppress. We observe that meaningful distinctions among on-policy trajectories can be distilled into natural-language process rubrics and used as process supervision signals. Based on this insight, we propose Online Rubric Self-Distillation(ORSD), which leverages self-derived process rubrics for dense token-level supervision, thereby augmenting outcome-centric reinforcement learning. Specifically, at each update, the current policy independently generates multiple rubric bundles from the same group of trajectories, extracting task-specific candidate process criteria from their cross-trajectory behavioral differences. A frozen verifier then evaluates all rubric criteria against each trajectory in the group, and selects one candidate bundle based on bundle-level alignment between rubric satisfaction and task success. ORSD treats the selected rubric bundle as privileged information during training, then compares the log-likelihoods of the original tokens with and without the rubric. The resulting likelihood gap provides token-specific process credit, which is added to the original trajectory-level GRPO advantage through additive advantage shaping. ORSD thus bridges trajectory-level outcome supervision and fine-grained process supervision, enabling token-level credit assignment using process signals. Extensive experiments on ALFWorld and WebShop show that ORSD accelerates policy improvement and boosts final task performance for both Qwen3-1.7B and Qwen3-4B under the same environment-interaction budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.