acceptodds
Under review as a conference paper at ICLR 2027

DriveAltas: A Cross-Country Benchmark for Grounded Driving-Rule Reasoning

Abstract

Vision–language models (VLMs) are becoming the perception-and-reasoning backbone of end-to-end driving systems, but the benchmarks used to assess them are built from one or two regions: they tell whether a model recognizes a scene, not whether it acts on the traffic rule that applies *where the scene was taken*. We introduce DriveAtlas, a cross-country, rule-grounded benchmark for *grounded driving-rule reasoning*, with 10,143 test and 26,539 training questions across eight countries and four tasks (perception, prediction, planning, and region). Each question is tied to a clause of the local traffic code, drawing on 5,962 unique rule texts, and is kept only if its correct answer changes under another country's rules; the country is never revealed to the model. We evaluate 16 open-weight VLMs (2B–9B parameters, three model generations) and report accuracy alongside a sampling-noise-corrected cross-country disparity . All 16 models show significant differences in driving performance across countries: ranges from 5.22 to 10.39, and its 95% interval excludes zero for every model. Scale does not remove these differences: a larger model of the same family is more accurate, but its country-dependent variation remains high, and in three of the four families where size buys accuracy, moves by under half a point. An error analysis of two models finds that naming the wrong country is rarely the failure (under 3% of sampled errors), whereas about one item in five fails on a regional rule gap. We therefore propose DriveOPD, an on-policy self-distillation algorithm whose teacher is the same model reading a jurisdiction-redacted traffic-rule handbook: on both backbones we train, it raises accuracy on all four tasks and lowers by about three points (6.91 → 3.87, 7.07 → 3.88). Used as the backbone of a vision–language–action policy, DriveOPD lowers open-loop L2 error and collision rate in both Boston and Singapore, and narrows the gap between the two cities by 19–26%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.