acceptodds
Under review as a conference paper at ICLR 2027

Not Every Step Is Progress: Event-Grounded Credit Reassignment for Robot Policy Learning

Abstract

Learning from deployment experience can improve robot policies beyond demonstrations. Sparse episode-level outcomes, however, obscure how credit should be assigned within a rollout. Successful rollouts may include locally ineffective behavior, whereas failed rollouts may still exhibit substantial productive execution. Even dense progress or value supervision can conflate effective task progress with elapsed interaction time. To address these issues, we introduce **Event-Grounded Credit Reassignment (ECR)**, which learns local transition effectiveness from episode outcomes, human interventions, and sparse annotations of failure onset. ECR models uncertainty about when ineffective execution begins and marginalizes over plausible onset locations to construct soft supervision for otherwise unlabeled transitions. The resulting estimates can then revise value targets for advantage-conditioned policy learning and reward signals for offline and online reinforcement learning. We instantiate these settings with RECAP, IQL, and RL Token, respectively. Across four simulation tasks and four real-world tasks, ECR improves the mean success rate over RECAP by 31.0 percentage points and achieves 2.78× the mean throughput of RECAP. In offline evaluation on the same task suite, ECR achieves a mean AUROC of 95.3% and a mean AP of 93.3% for distinguishing effective from ineffective transitions. Additional experiments with IQL and RL Token show consistent gains in offline and online reinforcement learning settings. Our results demonstrate that correcting local credit assignment improves policy learning from mixed-quality robot experience.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.