Driving Competence Is Not a Scalar: A Pareto Audit and Profile for End-to-End Driving
Abstract
A single number ranks end-to-end driving planners, and NAVSIM's PDMS, built to fix its predecessor, is that number. Each repair so far has adjusted weights inside the number and kept its type. We argue the type is wrong. Driving competence rests on attributes that genuinely trade off, so the object of scoring is a partial order rather than a quantity. On navtest, under a four-axis operationalization whose partition we also sweep, 62–71% of admissible trajectory pairs are Pareto-incomparable, and the per-scene competence poset has certified Dushnik–Miller order dimension ≥3 in ∼91% of scenes and dimension 1 in none. A scalar's own order is total, so on every such pair it must either collapse a real distinction or order it by an undeclared preference, and that contested set is exactly the incomparable set at any positive weighting. Known complaints follow. PDMS separates a human from a strong sensor-based planner by under 0.01 in 46.3% of scenes, its per-scene winner contradicts its aggregate ranking in up to 28.8% of scenes, and on a deployed pair that closed-loop testing separates by six points, the sign of its verdict turns on scorer implementation and weighting. The Driving Competence Profile (DCP) reports competence at its type. It returns an admissibility gate, four graded axes calibrated independently of the evaluated planners, and a leaderboard scalar carried with the profile and a map of which comparisons it arbitrates. Over 13 zoo planners plus current state-of-the-art methods, it ranks the human above a blind no-perception model in 83.4% of scenes against the official PDMS's 71.6%, and it flags evidence consistent with an external Goodhart step, a method whose published PDMS rose from 88.1 to 91.2 while its measured competence fell.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.