FORGE: An Execution-Feedback-Guided Framework for Benchmarking Generalist Robot Policies
Abstract
As Vision-Language-Action (VLA) models advance toward open-ended robotic manipulation, reliable evaluation of their performance across diverse capability dimensions becomes increasingly important. Yet comprehensive evaluation remains difficult due to the multimodal nature of VLA inputs and the cost of constructing executable simulation environments and reliable oracles at scale. Existing VLA benchmarks largely rely on manually specified or template-instantiated task-scene-oracle triplets constructed independently of model performance, resulting in limited input diversity, overlooked process failures, and only coarse-grained difficulty levels. In this paper, we introduce **FORGE**, an execution-feedback-guided framework that dynamically generates capability-specific mutations from validated seed tasks and scenes using observed VLA execution behavior, enabling iterative exploration of model failure modes. Using FORGE, we build **FORGE-Bench**, comprising 215 task-scene-oracle variants generated along validated mutation trajectories across five capability dimensions. Evaluation of five representative VLA models shows dimension-specific failures, terminal-process success gaps, and local performance reversals along shared mutation trajectories, providing a discriminative characterization of VLA capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.