acceptodds
Under review as a conference paper at ICLR 2027

CoFiDrive: Scene-Conditioned Registers and Coarse-to-Fine Trajectory Scoring for End-to-End Driving

Abstract

Vision-based driving planners must compress dense multi-camera observations while retaining the information needed to generate and compare motion alternatives. Road layout and traffic configuration provide scene context, whereas nearby trajectories can differ in their clearance from vehicles or lane boundaries. We present CoFiDrive, which combines scene-conditioned register refinement with coarse-to-fine trajectory scoring. Four camera views contribute 16 learned registers each, forming a compact 64-token scene memory. A semantic token compressor (TC) refines these 64 tokens through scene-conditioned soft expert mixing, preserving the memory size for trajectory generation and scoring. The Cross-Compression Scorer (CCS) combines the refined scene features with spatially pooled patch features through candidate-dependent gated attention, giving each proposal access to coarse context and local spatial detail. During training, control-space augmentation (CSA) derives acceleration and yaw-rate sequences from PDM-ranked proposals, perturbs these controls, and uses bicycle-model rollouts to create additional candidates. The PDM evaluator supplies component-score targets for the same CCS on the original and augmented candidates. CSA broadens CCS supervision without adding model parameters. At inference, CCS ranks only the original 64 proposals, with no augmentation or PDM target computation. After 10 training epochs, CoFiDrive achieves 92.4 and 92.9 PDMS on NAVSIM-v1 using train85k and trainval103k, respectively, and 48.0 EPDMS on NAVSIM-v2 NavHard.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.