QuoVLA: Quotient Space for Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models adapt pretrained visual-language features for robot control through action supervision. Action supervision can guide task-specific feature learning and the suppression of action-irrelevant variation. To limit this variation, we introduce QuoVLA, which combines joint feature learning with a quantized prefix adapter. The adapter processes visual features and language embeddings before a shared language backbone and action expert, using flow matching and relative temporal-complexity regularization. This regularization uses a stop-gradient raw-prefix reference to discourage excess temporal variation in the flow predictions of the learned policy. We interpret this policy through an action quotient that groups representations with the same chosen optimal action law. For this fixed action law, a conditional characterization defines an ideal compression target without assuming sufficient pretrained features or guaranteeing that training attains the target. The target motivates paired representation diagnostics alongside task success evaluation. Task success improves over the primary baseline on four simulation benchmarks and four real-robot tasks, while the diagnostics support selective compression under the evaluated perturbations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.