Model-Agnostic Contrastive Adaptation for Slow-Fast End-to-End Autonomous Driving
Abstract
The dual-system architecture combining fast-reactive end-to-end (E2E) models with slow-reasoning Vision Language Models (VLMs) has emerged as a promising solution for autonomous driving in complex scenarios. However, the rapid iteration of E2E models demands a flexibly decoupled, model-agnostic framework. While injecting latent features from VLMs into E2E models preserves dense semantic guidance and inherently offers a higher theoretical upper bound for performance, naive implementations often fail, forcing researchers to adopt intricate, deeply coupled fusion structures. In this paper, we reveal that this plugin failure is not a structural deficiency but an optimization trap. Due to cross-model shortcut learning and gradient vanishing, pre-trained E2E models instinctively treat injected semantic priors as perturbation noise. To address this, we propose MACA, a novel model-agnostic cognitive plugin. MACA bypasses high-dimensional feature injection by extracting concise semantic intentions via a two-stage Sequential Reasoning-to-Decision Training (SRDT) strategy. To break optimization stagnation, we introduce Relative Advantage Preference Optimization (RAPO), which evaluates the VLM-augmented model against a frozen E2E baseline to generate dynamic comparative signals, compelling the network to deeply internalize VLM cognition. Extensive evaluations demonstrate MACA's generality across diverse state-of-the-art E2E architectures, including DiffusionDrive, SparseDrive, MomAD, and GoalFlow. For instance, integrating MACA into DiffusionDrive reduces the average L2 error from 0.57 m to 0.51 m and the collision rate from 0.08% to 0.05% on nuScenes, while improving the PDMS of DiffusionDrive and GoalFlow from 87.5 and 89.2 to 88.2 and 90.0, respectively, on the NAVSIM benchmark. Our approach ensures driving stability in regular scenarios while significantly boosting decision-making success in corner cases, comprehensively outperforming existing baselines. Code and models will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.