acceptodds
Under review as a conference paper at ICLR 2027

Learning When and What Data to Use for Vision-Language-Action Models

Abstract

Vision-language-action (VLA) policies increasingly use heterogeneous observations, but the relevance of each data source can change during task execution. We study which already encoded sources should continue to participate in action prediction, conditioned on the context formed by the policy. We propose DataRouter, a decision-conditioned source-routing framework for pretrained VLA policies. DataRouter combines source-level granularity with decision-state conditioning: native source representations and intermediate VLA states are mapped into a shared space to estimate context-dependent source utilities. One trained utility estimator supports fixed-budget, threshold, adaptive-budget, and cost-aware inference rules. Only lightweight routing components are trained from demonstrated actions, while the pretrained VLA parameters remain frozen. Continuous gates provide a training surrogate, while hard routing omits unselected source tokens from subsequent action prediction. Across five VLA configurations, DataRouter yields average relative improvements of 8.58% on LIBERO-Plus and 11.16% on CALVIN, while retaining only 45% and 44% of the available source groups for subsequent action prediction, respectively. These results support decision-dependent utilization of encoded sources in the evaluated settings, while preserving their earlier influence and encoding cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.