acceptodds
Under review as a conference paper at ICLR 2027

GTPO: Group Teaching Policy Optimization via Temporal Optimal Transport in VLA Reinforcement Learning

Abstract

Reinforcement learning can improve vision-language-action policies beyond their demonstrations, but costly robot interactions and sparse terminal feedback make this process inefficient. Reliable dense supervision could extract more useful information from each interaction, yet learned reward models often require substantial additional data construction and training. Can the policy's own successful and failed rollouts provide this supervision? We propose Group Teaching Policy Optimization (GTPO), which turns successful rollouts into teachers for failed attempts at the same task from similar initial states. Building on Group Relative Policy Optimization (GRPO) and drawing inspiration from inverse reinforcement learning, GTPO uses trajectory comparison to refine otherwise uniform trajectory-level advantages. Each failed rollout is paired with its nearest successful in-group teacher using the policy's visual features. Temporally masked optimal transport aligns their trajectories and produces normalized action-chunk weights that redistribute the original GRPO advantage while retaining terminal rewards. The teachers evolve with the policy, requiring no additional expert reference trajectories or separately trained reward model for teaching. We evaluate GTPO on LIBERO, RoboTwin, and five challenging real-world tasks requiring fine manipulation. Across four LIBERO suites, GTPO achieves 98.65% average success with one demonstration per task and 99.70% with the full demonstration set, requiring at least 22.2% fewer aggregate training steps to reach 95% success than matched GRPO. On the five real-world tasks, GTPO achieves 80–100% success after three on-policy rounds, exceeding GRPO by 10–35 percentage points. Project website: https://gtpo-project.github.io/gtpo_page/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.