acceptodds
Under review as a conference paper at ICLR 2027

Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models

Abstract

Long-horizon robotic manipulation requires vision-language-action (VLA) models to track scene states and their evolution beyond the current observation. However, simply conditioning policies on observation history does not guarantee that the history is effectively utilized: action supervision constrains what the policy should do, but only indirectly constrains what its history representations should retain. To resolve this, we present Temporal Forcing, a 4D representation alignment framework that explicitly supervises latent temporal states and their transitions. Specifically, we first introduce a history pathway that compresses past observations into compact latent tokens. We then align these tokens and current-frame features with geometric targets from a pretrained 4D foundation model, providing direct supervision at both the state and transition levels. The 4D foundation model and alignment heads are used only for training-time supervision. Temporal Forcing improves average success from 96.6% to 98.8% on LIBERO, with the largest gain on LIBERO-Long (93.8% to 97.2%), and from 53.5% to 62.8% across twelve RoboTwin 2.0 tasks. Furthermore, Temporal Forcing increases full-task success from 20.0% to 43.3% on a physical multi-stage hidden-placement task. Controlled experiments show that 4D representation alignment is crucial for making observation history beneficial to the model. Code will be publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.