acceptodds
Under review as a conference paper at ICLR 2027

Pelican-VLA 0.5: Inducing Transferable Manipulation-Centric Attention via Bottleneck Tokens

Abstract

Generalization across objects, scenes, tasks, and embodiments remains a central challenge for Vision-Language-Action (VLA) models, which still rely heavily on task- and environment-specific robot data. We introduce Pelican-VLA 0.5, a unified model for vision-language understanding, future-frame prediction, and action generation. At its core is a compact set of learnable Bottleneck Tokens (BotTokens), through which information is routed to the action pathway. This bottlenecked perception-to-action routing gives rise to attention-level generalization: without object annotations, segmentation masks, or explicit attention supervision, the model develops manipulation-centric attention during pre-training. Even before task-specific fine-tuning, its action pathway can localize task-relevant objects and contact regions across unseen scenes and robot embodiments. After fine-tuning, Pelican-VLA 0.5 achieves 98.6% success on LIBERO, 91.4% on RoboTwin Clean and 91.0% on RoboTwin Randomized. Across three real-robot embodiments, it performs comparably to under the same evaluation setting. Moreover, we find that its attention patterns remain highly consistent before and after fine-tuning, providing further evidence that the manipulation-centric attention acquired during pre-training is transferable. At the same time, the attention-to-action gap between correctly attending to manipulation-relevant regions and successfully executing actions remains a key obstacle to developing broadly generalizable robot policies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.