DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?
Abstract
Frontier models match expert humans across many digital benchmarks, yet real-world driving, a skill any new driver has, remains untested for them. We present DrivingBench, the first benchmark where general-purpose vision-language models must drive a real car. Through three tools, the models see camera frames from a Toyota Corolla and directly command its steering and speed around a parking lot cone course. The car may continue moving while the model thinks and new commands replace the currently running one, so inference latency is part of the task, testing the models' abilities to observe, act, monitor, recover, and complete a long-horizon objective under such constraints. We benchmark GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol, and Grok 4.6 in vendor-native harnesses (Codex, Claude Code, Cursor) with up to three attempts each in one conversation; Astra is the only model to finish the course, on its second attempt, with no other attempt passing 50% of the course. Two of the four models improved materially across attempts with retained context. We release our harness, prompts, course map, and traces with video and telemetry for reproducibility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.