acceptodds
Under review as a conference paper at ICLR 2027

ESCAPO: Entropy-guided Segment-level Credit Assignment for Policy Optimization in VLA Post-Training

Abstract

Long-horizon vision-language-action (VLA) policies are increasingly post-trained with on-policy reinforcement learning using sparse terminal rewards. Critic-free methods such as Group Relative Policy Optimization (GRPO) derive a trajectory-level advantage from these outcomes, providing only coarse temporal credit: a failed rollout may contain useful intermediate behavior, yet every action receives the same advantage. We find that the policy itself provides a finer signal. Human-annotated rollout analysis shows that, after supervised fine-tuning (SFT), locally successful actions tend to have lower policy entropy than failed actions within the same trajectory, and that this separation becomes stronger during on-policy RL post-training. Based on this observation, we propose Entropy-guided Segment-level Credit Assignment for Policy Optimization (), which uses policy entropy to refine GRPO advantages without a critic, dense reward, teacher policy, or action-quality labels for training. In failed rollouts, reliable low-entropy segments receive increased credit while high-entropy segments are penalized; in successful rollouts, uncertain actions supported by successful outcomes receive additional positive credit. On LIBERO-10 and LIBERO-Plus-10, improves long-horizon learning over GRPO from both SFT and shared RL warm-up initializations. From SFT, it raises late-stage success from to and from to , respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.