acceptodds
Under review as a conference paper at ICLR 2027

Structured Latent Reasoning for Vision-Language-Action Models

Abstract

Vision-language-action (VLA) models can benefit from intermediate reasoning, but explicit chain-of-thought reasoning incurs substantial autoregressive decoding overhead in closed-loop control. Existing latent reasoning methods reduce this cost through continuous latent representations, yet leave the structured reasoning roles of explicit traces implicit. We propose Structured Latent Reasoning (SLR), which organizes latent representations into semantically distinct reasoning units corresponding to different forms of intermediate reasoning and supervises them through auxiliary text decoding. Because the relevance of these units varies across manipulation transitions, we further introduce transition-adaptive learning, which reweights their supervision according to relevance signals from the action policy. At inference, SLR operates directly on the structured latent units, avoiding autoregressive reasoning generation. SLR achieves 73.7% average success on SIMPLER, improving over LaRA-VLA by 4.9 points, and reaches 79.0% average task completion on real-robot tasks while maintaining gains under distribution shift. SLR runs at 341 ms per control step, combining the structured semantics of explicit reasoning with the inference efficiency of latent reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.