Beyond Fixed Benchmarks: Measuring Data-Scaling Behavior of Autonomous-Driving Planners on the Physical AI Autonomous Vehicles Dataset
Abstract
Public autonomous-driving benchmarks typically compare planners at a single training-data size. This identifies the top-ranked planner at that size but not how performance or rankings change with more data. Using the NVIDIA Physical AI Autonomous Vehicles Dataset (PhysicalAI), we train five planners on common nested subsets ranging from 1 to 64 hours. We evaluate the selected checkpoints on a fixed, training-disjoint PhysicalAI set using three-second average and final displacement errors (ADE3 and FDE3), and on NAVSIM v2 navtest using the closed-loop-relevant Extended Predictive Driver Model Score (EPDMS). The scaling curves reveal three behaviors hidden by a single benchmark point: the top-ranked planner according to ADE3 changes with training-data size; planners achieve different gains after the same data doubling; and lower open-loop error does not always lead to higher EPDMS. Uni-World VLA has the highest post-adaptation EPDMS at every training-data size, yet its score changes by only 0.268 points from 1 to 64 hours. These results separate absolute performance from data-scaling behavior and show why comparisons among public planners should report scaling curves rather than one fixed-size result.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.