acceptodds
Under review as a conference paper at ICLR 2027

Segmentation-Guided Semantic Distillation For Robot Manipulation

Abstract

Pretrained Vision Transformer (ViT) features may contain high-response background tokens and exhibit spatial-semantic misalignment, which can weaken the spatial perception of Vision-Language-Action (VLA) models and thereby impede spatially precise robotic control. Action-supervised adaptation does not reliably remove these artifacts. Therefore, we propose Segmentation-Guided Semantic Distillation (\method), which transfers dense features from a CNN segmentation teacher trained on robotic manipulation data directly into the VLA visual backbone. During policy fine-tuning, three auxiliary losses, namely spatially corresponding feature matching, directional alignment, and magnitude alignment, are jointly optimized with the action prediction loss. On five RoboTwin tasks and three real-world robot-arm tasks, improves success rates over frozen-backbone baselines by 17.5 and 18.7 percentage points, respectively. These results show that task-matched spatial distillation improves both foreground-aligned representations and manipulation success. Code and models are available at: https://anonymous.4open.science/r/SGSD-Segmentation-Guided-Semantic-Distillation-0760.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.