acceptodds
Under review as a conference paper at ICLR 2027

TARA: Temporal Action-Reasoning Alignment for Steerable Vision-Language-Action Policies

Abstract

Recent research increasingly explores generalist Vision-Language-Action (VLA) policies that leverage large-scale vision-language pretraining to generalize from only a few demonstrations. A promising direction is Chain-of-Thought reasoning, which enables models to reason step-by-step about the environment state and next subgoals before generating low-level actions. However, current VLAs tend to ignore intermediate thoughts when executing actions as they over-rely on vision. Prior work has tried to address this with information-dense intermediate representations, but such approaches struggle in low-data settings. We therefore propose a data-efficient framework TARA (Temporal Action-Reasoning Alignment) that improves policy steerability through precise temporal alignment between reasoning and actions. We heuristically decompose demonstrations into meaningful subgoals using implicit visual and behavioral information and train flow-matching policies to predict the actions required for each subgoal. We show that this temporal alignment increases attention to reasoning traces while preserving learned fine-grained behavioral knowledge. Moreover, aligned policies learn latent representations that better support compositional generalization and demonstrate robustness on LIBERO-Plus (+3.7%), generalizability through improved policy steerability on specifically constructed benchmark (+4.8%) across different subgoal representations, and significant improvements on a real-world Franka robot.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.