acceptodds
Under review as a conference paper at ICLR 2027

CSDR: Control-Structured Dispersive Regularization for Vision–Language–Action Models

Abstract

Vision–language–action (VLA) models transfer semantic understanding to robot control through pretrained multimodal backbones. However, representations learned from image–text pretraining may not adequately distinguish different control requirements. To address this limitation, we propose CSDR, a control-structured representation regularization method applied only during training. Its central idea is that representations need not reproduce absolute distances between numerical actions; instead, under the same instruction, samples with similar control requirements should have closer representations than those with substantially different requirements. We identify these relationships using valid action sequences from existing demonstrations and apply a ranking loss that prioritizes samples with larger action prediction errors. Each ordering constraint becomes inactive once its prescribed margin is satisfied. CSDR introduces no teacher model, additional projection head, or inference module. Experiments with four VLA baselines on LIBERO and BridgeData V2 demonstrate that CSDR effectively improves policy performance by refining internal representations. Real-world robot experiments further corroborate its effectiveness. These results establish CSDR as a simple and transferable approach to improving VLA performance. The code is available at https://anonymous.4open.science/r/CSDR-VLA-9A7F/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.