acceptodds
Under review as a conference paper at ICLR 2027

-sight: Unifying Vision-Language-Action and Latent World Modeling

Abstract

We introduce -sight, a unified robot foundation model that brings vision-language-action modeling and world modeling together through feature-space future prediction. Recent video-based world action models (WAMs) adapt pretrained video generators for robot control, combining visual dynamics modeling with action learning at substantial computational cost. -sight brings future prediction into the native visual representation space of a pretrained vision-language model (VLM), allowing the policy to perceive the present and predict the future within a shared visual foundation. We couple a dedicated foresight expert and a continuous action expert to the VLM through a mixture-of-transformers architecture. The foresight expert denoises future spatial visual tokens, while the action expert attends to the prediction stream to generate actions. This design supports joint learning from robot trajectories and human egocentric video without requiring pixel reconstruction or a pretrained video generator. During inference, -sight efficiently generates future visual features and continuous actions through joint denoising. Evaluated on a real humanoid robot, -sight achieves a 62% average success rate across five manipulation tasks, outperforming VLA and WAM baselines by a large margin. Furthermore, it maintains real-time inference speed of 110ms per chunk without the latency bottleneck of pixel-space generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.