acceptodds
Under review as a conference paper at ICLR 2027

AutoPrune: Accelerating End-to-End Driving Training via Gradient-Preserving Adaptive Data Pruning

Abstract

Unified end-to-end architectures have reshaped autonomous driving, but their reliance on massive datasets leads to weeks-long training, severely hindering model iteration. A natural way to accelerate training is to remove simple or redundant data that provide limited learning value. However, existing data reduction methods for end-to-end driving often rely on static data features and hand-crafted selection rules, facing two fundamental limitations: (1) Gradient distortion. Pruning low-value data can alter the gradient composition and degrade model performance; (2) Training-dynamics misalignment. Evolving data values make fixed pruning decisions stale, while fewer iterations can misalign the optimization schedule with full-data progress. Therefore, we present AutoPrune, which reduces training-data usage to accelerate end-to-end driving training while mitigating pruning-induced performance degradation. First, Gradient-Preserving Data Pruning is introduced to compensate for the missing training gradients caused by pruning low-loss samples, so that the corrected aggregate gradient matches its full-data counterpart in expectation. Second, Training-Dynamics Adaptation uses loss-adaptive sample gating to update pruning decisions from the latest available loss records and pruning-aware learning-rate alignment to follow the corresponding full-data training progress. Experiments across multiple end-to-end driving models show that AutoPrune retains performance close to the corresponding full-data references at a 40% overall pruning ratio. Comparative and ablation studies further support the effectiveness of the proposed components.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.