LIBERO-Pro: A Benchmark for Evaluating the Robustness of Embodied Foundation Models
Abstract
LIBERO has become a de facto standard benchmark for evaluating embodied intelligence. However, high success rates on LIBERO do not necessarily reflect robust embodied capabilities, as models may achieve strong performance by exploiting benchmark-specific patterns. In this work, we introduce LIBERO-Pro, a comprehensive benchmark for systematically evaluating the robustness of embodied foundation models. LIBERO-Pro defines 22 static perturbations and 20 runtime interventions, resulting in 1,648 validated perturbation variants. These perturbations cover changes in scene configurations, object properties, visual observations, robot states, temporal conditions, task instructions, and runtime execution states. Each perturbation is carefully designed and undergoes six rounds of human verification to ensure physical plausibility and task validity. We evaluate 11 representative robot policies spanning vision-language-action models, world-action models, and robustness-oriented policies. Our results reveal substantial robustness gaps across all evaluated paradigms: under runtime distribution shifts, the evaluated models suffer an average performance drop of 25.9 percentage points, with individual drops reaching 32.5 points. These findings demonstrate that strong nominal benchmark performance does not necessarily translate into robustness under distribution shifts, highlighting the need for robustness-oriented evaluation of embodied foundation models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.