Which Synthetic Driving Data Are Worth Training On? Dynamic Data Valuation Based on Their Effects on Model Training
Abstract
Generative models can now produce synthetic driving data at scale, providing a promising way to improve the performance of end-to-end driving models beyond costly real-world data collection. However, not all synthetic data are equally valuable for training, making it important to accurately identify which data are worth training on. Existing methods typically value synthetic data using their properties, such as visual realism and temporal consistency, thereby treating data value as an invariant property of the data. This overlooks two key issues: (i) these data properties do not directly reveal whether synthetic data actually improve model performance; and (ii) the value of synthetic data is not fixed, but varies across training stages and can be sensitive to the data used for valuation. To address these issues, we propose DriveVal, which evaluates synthetic driving data based on their training effects on end-to-end driving model performance and uses the resulting values for data selection. With a small set of real driving data as references, data value is quantified by the dot product between the real-data gradient and the synthetic-data gradient under evaluation. In addition, DriveVal periodically re-evaluates data throughout training while accounting for the current optimizer state at each stage, allowing the valuation to adapt as training progresses. Finally, DriveVal penalizes variations in data value across different real references and synthetic-data combinations, preventing the valuation from being dominated by a particular data choice. Across five end-to-end driving models on NAVSIM, DriveVal selects synthetic training data matched in size to navtrain throughout training. The resulting models achieve 95.2% of the mean PDMS of their corresponding real-data baselines, while improving mean PDMS by 13.6% over size-matched random selection and by 9.1% over training on the full synthetic set. Further experiments demonstrate that DriveVal remains effective across different generation sources and visual conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.