acceptodds
Under review as a conference paper at ICLR 2027

CASA: Cluster-Adaptive Scaling for Hierarchical Data Selection in End-to-End Autonomous Driving

Abstract

As autonomous driving datasets continue to grow, efficiently constructing training sets from large-scale data is becoming increasingly important. Effective selection therefore requires considering not only which individual samples are valuable, but also how the training budget should be distributed across different scene types. We introduce CASA, a hierarchical data-selection framework that jointly addresses these two levels. CASA first organizes driving data into behaviorally meaningful clusters by coupling semantic context with future motion, and performs task-aware ranking within each cluster. At the cluster level, CASA learns a structure-aware scaling model with shared parameters across clusters while accounting for shared coverage between related clusters, allowing a small number of probes to guide marginal-utility-based budget allocation. By integrating sample-level selection with structure-aware cluster allocation, CASA provides a unified approach to constructing training sets under varying budgets. On NAVSIM with DiffusionDrive, CASA learns effective scaling relationships from low-budget probes and maintains the highest mean Predictive Driver Model Score (PDMS) among five representative data-selection baselines even when the training budget exceeds four times the largest probe size. This sustained advantage demonstrates CASA’s effectiveness in improving data efficiency for end-to-end driving.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.