From Sparse to Dense: Diagnosing Task-Equivalent Language Robustness in Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models are commonly trained and evaluated with concise task commands, leaving open whether learned control remains accessible when the same executable intent is expressed through richer, task-equivalent language. An audit of LIBERO and SimplerEnv finds that their native instructions concentrate on direct, low-context expressions. To study this uncovered regime, we construct a three-level task-equivalent benchmark with controlled dense language and evaluate models under matched rollouts that vary only the language realization. These evaluations reveal a substantial Sparse–Dense Gap whose magnitude depends strongly on the model and task. Further diagnostics show that paired sparse–dense action training yields limited dense gains and that action fine-tuning reduces inherited visual-language capability. Guided by these findings, STAR, a Sparse-to-dense Task Alignment Regularization framework, uses a lineage-matched pre-action VLM and the source VLA as complementary representation and action anchors. Across LIBERO, STAR raises dense macro success from 7.2% to 24.3% while retaining 69.7% sparse success, and it improves the corresponding sparse–dense balance in SimplerEnv, demonstrating that task-equivalent dense-language access is a distinct robustness problem that can be systematically measured and improved.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.