acceptodds
Under review as a conference paper at ICLR 2027

QuoVLA: Quotient Space for Vision-Language-Action Models

Abstract

Vision-Language-Action (VLA) models adapt pretrained visual-language features for robot control through action supervision. Action supervision can guide task-specific feature learning and the suppression of action-irrelevant variation. To limit this variation, we introduce QuoVLA, which combines joint feature learning with a quantized prefix adapter. The adapter processes visual features and language embeddings before a shared language backbone and action expert, using flow matching and relative temporal-complexity regularization. This regularization uses a stop-gradient raw-prefix reference to discourage excess temporal variation in the flow predictions of the learned policy. We interpret this policy through an action quotient that groups representations with the same chosen optimal action law. For this fixed action law, a conditional characterization defines an ideal compression target without assuming sufficient pretrained features or guaranteeing that training attains the target. The target motivates paired representation diagnostics alongside task success evaluation. Task success improves over the primary baseline on four simulation benchmarks and four real-robot tasks, while the diagnostics support selective compression under the evaluated perturbations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.