acceptodds
Under review as a conference paper at ICLR 2027

HRAE: One-Step Action Generation for Flow-Matching VLAs via Hierarchical Recurrent Action Expert

Abstract

Diffusion and flow-based vision-language-action (VLA) policies achieve strong robotic control through repeated model evaluations to progressively refine actions, but such multi-step inference introduce substantial latency. Existing one-step approaches reduce this cost by compressing the external generative process, often through distillation, consistency objectives, or specialized training targets, which can require pretrained teachers or additional trajectory-level training machinery. However, reducing external sampling steps also removes the iterative computation performed during generation. Recent evidence suggests that such iterative computation, when appropriately supervised, is an important inductive bias for generative robotic control. This raises a key question: can we retain iterative computation while avoiding repeated external flow evaluations? We address this by moving iterative computation inside each action-expert evaluation. We introduce the **Hierarchical Recurrent Action Expert (HRAE)**, which couples a fast low-level latent state for recurrent action refinement with a persistent high-level latent state that integrates information across refinements and guides subsequent computation. By reusing the same low-level and high-level modules across hierarchical recurrent cycles, HRAE achieves substantial effective computational depth without proportionally increasing model size. We further pair HRAE with clean-action -prediction, which aligns the prediction target with its internally refined action representation while remaining compatible with multi-step flow integration. Across tasks in LIBERO, LIBERO Plus, RoboTwin, and real-world manipulation, HRAE achieves strong in-distribution performance and robustness under distribution shifts with only one external flow evaluations. In real-world tasks, one-step HRAE outperforms a 10-step GR00T baseline while reducing action-expert latency from 44.73 ms to 6.85 ms. These results suggest that iterative computation need not be tied to repeated external flow integration and can instead be effectively realized within the action expert for efficient generative robotic control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.