FlatVLA: Robust Vision-Language-Action Model with Sharpness-Aware Minimization
Abstract
Vision-Language-Action (VLA) models integrate visual observations and language instructions to generate actions, yet remain vulnerable to multimodal distribution shifts and errors arising during sequential execution. Existing robustness methods target predefined variations, making it difficult to cover a broad range of variations across VLA inputs and actions. To address this challenge, we propose FlatVLA, a flatness-aware training framework that adapts Sharpness-Aware Minimization (SAM) to the sequential structure of VLA action prediction. FlatVLA applies larger weights to positions near the beginning of each action chunk, as errors at these positions can affect a larger portion of the remaining prediction. Our theoretical analysis connects local smoothness to policy sensitivity and compounding error, motivating this sequential weighting. Across recent models, FlatVLA shows improvements on performance and robustness across diverse evaluation factors and settings, with further analysis of model behaviors. These gains also extend from simulation to real-world tasks, as demonstrated via hardware experiments. Overall, FlatVLA optimizes parameter space smoothness while accounting for sequential error accumulation in action prediction, improving robustness without explicitly tailoring training to specific deployment conditions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.