CSDR: Control-Structured Dispersive Regularization for Vision–Language–Action Models
Abstract
Vision–language–action (VLA) models transfer semantic understanding to robot control through pretrained multimodal backbones. However, representations learned from image–text pretraining may not adequately distinguish different control requirements. To address this limitation, we propose CSDR, a control-structured representation regularization method applied only during training. Its central idea is that representations need not reproduce absolute distances between numerical actions; instead, under the same instruction, samples with similar control requirements should have closer representations than those with substantially different requirements. We identify these relationships using valid action sequences from existing demonstrations and apply a ranking loss that prioritizes samples with larger action prediction errors. Each ordering constraint becomes inactive once its prescribed margin is satisfied. CSDR introduces no teacher model, additional projection head, or inference module. Experiments with four VLA baselines on LIBERO and BridgeData V2 demonstrate that CSDR effectively improves policy performance by refining internal representations. Real-world robot experiments further corroborate its effectiveness. These results establish CSDR as a simple and transferable approach to improving VLA performance. The code is available at https://anonymous.4open.science/r/CSDR-VLA-9A7F/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.