acceptodds
Under review as a conference paper at ICLR 2027

WAM-Bench: A Closed-Loop Benchmark for Open-Ended Instruction Following in Autonomous Driving

Abstract

Recently, Vision-Language-Action (VLA) models have shown great potential in advancing end-to-end autonomous driving (E2E-AD). However, existing E2E-AD benchmarks typically provide only two predefined instructions (go straight and turn), leaving the ability of VLA models to follow human instructions largely unexplored. To address this limitation, we introduce WAM-Bench, a real-world closed-loop benchmark for evaluating open-ended instruction following in E2E-AD. It contains 37 hours of driving videos paired with human-reviewed open-ended instructions and covers 13 action classes. We further establish a comprehensive evaluation framework with representative VLA baselines, and propose an instruction-following metric (IF) for closed-loop evaluation. Extensive experiments across diverse settings, including the standard test split (with various action classes), long-horizon execution, changed instructions, and OOD instructions, reveal that open-ended instruction following remains challenging for current VLA models. The code is available at -.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.