UATL: Uncertainty-Aware Temporal Learning for Open-Loop Vision-Language-Action Control
Abstract
Vision-language-action (VLA) models map image observations and language instructions directly to robot actions, but the need for a full forward pass for each action prediction hinders the action rates required for control. Action chunk addresses this by predicting and executing multiple actions open-loop from a single policy query, reducing policy query frequency and increasing VLA action throughput. However, existing action chunk VLA methods based on deterministic continuous regression treat all timesteps and action dimensions within a chunk uniformly, capturing neither differences in predictive confidence nor temporal variation between consecutive actions. Consequently, local prediction errors accumulate during open-loop execution, degrading policy stability in fine-grained manipulation and long-horizon sequential tasks. To address this issue, we propose an Uncertainty-Aware Temporal Learning (UATL) framework for open-loop vision-language-action control. Specifically, we introduce Heteroscedastic Action-chunk Learning (HAL) to model continuous action prediction as a conditional Laplace distribution and jointly learn the action mean together with per-timestep, per-dimension scale parameters via negative log-likelihood, thereby explicitly characterizing aleatoric uncertainty in action prediction. Furthermore, we design a second-order Uncertainty-Weighted Temporal Consistency Regularizer (UW-TCR) that uses the reciprocal of the predicted scale as a per-timestep, per-dimension regularization weight, suppressing irregular jitter during open-loop execution while preserving the rapid action changes required at contact and direction reversals. Finally, evaluations on LIBERO, LIBERO-Plus, SimplerEnv, and CALVIN show improved performance in fine-grained manipulation, distribution-shift robustness, and long-horizon control, with negligible additional inference overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.