acceptodds
Under review as a conference paper at ICLR 2027

TAGER: Temporal Action Geometry for Robust Robot Manipulation

Abstract

Vision-language-action (VLA) models inherit rich visual-semantic knowledge from pretrained vision-language models (VLMs), yet remain vulnerable to distribution shifts in visual observations, language instructions, and manipulation conditions. A key limitation is that pretrained VLMs lack explicit geometric awareness and spatial relationships required for manipulation. Recent methods mitigate this limitation by aligning VLA representations with geometric teacher features. However, directly injecting such supervision into shared visual representations may interfere with the visual-semantic priors inherited from pretraining. Moreover, static spatial representations fail to capture the action-conditioned evolution of visual observations and robot states, thereby restricting the policy's capacity to adapt subsequent actions under perturbations. To address these challenges, we introduce TAGER, a Temporal Action Geometry framework for robust Embodied Robot manipulation that couples Sparse Spatial Query Distillation with an Action Response Residual Flow. To enhance spatial awareness while while mitigating interference with pretrained visual-semantic priors, sparse spatial query distillation uses dual independent query transformers to distill dense geometric and semantic knowledge from frozen VGGT and SAM3 teachers into compact tokens. These tokens provide complementary world-grounded context to the action expert without directly imposing teacher-feature matching constraints on shared VLA representations. To support adaptive action generation under perturbations, the action response residual flow relates recent executed actions to subsequent changes in robot states and spatial tokens representations, providing evidence about the current visual-motor relationship. The resulting execution context drives a residual flow that adaptively refines subsequent action generation. Experiments on LIBERO and VLABench, together with zero-shot evaluations on LIBERO-Plus and LIBERO-PRO, demonstrate that TAGER achieves outstanding manipulation performance while generalizing robustly across diverse perturbations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.