Thinking in Trajectories: Structured Exploration for Vision-Language-Action Reinforcement Learning
Abstract
Reinforcement-learning post-training has recently improved vision-language-action (VLA) models beyond supervised imitation, yet most existing methods explore by sampling low-level actions or action chunks. In a high-dimensional continuous action space, such perturbations mostly yield local control variations: changing the high-level intent of a long-horizon behavior requires coordinated deviations over many steps, which random action noise discovers only inefficiently. We propose TCoT-RL, which formulates trajectory reasoning as a directly optimizable exploration policy for VLA post-training: a trajectory policy autoregressively samples a variable-length sequence of end-effector waypoints representing a high- level motion plan, while a flow-based action policy generates continuous controls conditioned on the sampled trajectory. This hierarchical factorization enables structured exploration over alternative behavioral intents rather than only perturbing low-level actions. To directly optimize both levels of the hierarchy, TCoT-RL performs joint trajectory–action reinforcement learning with level-specific credit assignment. We introduce dense trajectory feedback based on task progress and stage-aware endpoint grounding, and optimize the trajectory and action policies with separate group-relative advantages and clipped GRPO objectives. This allows trajectory reasoning to evolve from a supervised intermediate representation into an active exploration policy through online interaction. Experiments show that TCoT-RL consistently improves over supervised trajectory-conditioned policies and matched action-only RL baselines. On LIBERO-Plus, TCoT-RL achieves a state-of-the-art success rate of 89.6%, outperforming the strongest competing RL baseline by 1.8 percentage points. Paired rollout analysis further shows that trajectory-level exploration expands end-effector state coverage and visitation entropy and produces greater behavioral diversity. These results demonstrate that explicitly optimizing the distribution over high-level motion hypotheses provides an effective structured exploration mechanism for VLA reinforcement learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.