acceptodds
Under review as a conference paper at ICLR 2027

Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer

Abstract

Vision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision–Language Models (VLMs). However, such VLM backbones offer insufficient physical dynamics priors, which limits the generalization capabilities of robot policies. Recent efforts therefore integrate video-generation World Models (WMs) into robot policies through various strategies, using predictive dynamics to facilitate action generation. Despite these advances, harnessing semantic understanding and dynamics prediction as complementary guidance for action generation remains challenging. In this paper, we introduce ACT, a simple yet effective Action-Centric Tri-Stream Transformer that fuses semantic and dynamics information into control action while preserving the distinct roles of context streams. Specifically, ACT enables the dedicated action expert to access VLM and WM representations through layerwise attention, with each backbone attending only within its own streams. This straightforward interaction design maintains independent forward propagation in the context streams while allowing both backbones to be updated through control supervision. Experiments on both simulated and real-world robotic manipulation benchmarks show that the proposed ACT yields results superior to its counterparts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.