acceptodds
Under review as a conference paper at ICLR 2027

GF-VLA: Dual Geometry Forcing for Vision-Language-Action Models

Abstract

Vision-language-action (VLA) models generate robot actions from visual observations and language instructions, but recognizing a manipulation goal does not determine how the end effector should move through the surrounding scene. We introduce **Geometry Forcing for VLA (GF-VLA)**, which brings both motion and scene geometry together on a shared action-conditioning pathway. Motion geometry forcing supervises time-indexed motion latents with positions and successive displacements from a sparse, temporally ordered 3D end-effector (EEF) trace within an action chunk. Scene geometry forcing reconstructs future depth at the chunk endpoint, guiding the scene latents read by motion queries. These motion latents, together with native vision-language features, condition a flow-matching Action Expert, supporting embodied reasoning that jointly considers 3D motion and future scene geometry during action generation. GF-VLA achieves 98.40% mean success on LIBERO and 83.8% overall zero-shot success on LIBERO-Plus; on RoboCasa-GR1, it improves success by 6.3 pp over a strong VLA baseline. Across three real-world tasks, GF-VLA achieves 73.3% average success under standard conditions, exceeding the baseline by 12.2 pp. These results support the benefit of coupling 3D motion trajectories with scene geometry at the action-chunk endpoint for action generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.