acceptodds
Under review as a conference paper at ICLR 2027

Learning How to Drive at Test Time

Abstract

Driving policies are often deployed with verifiers, i.e. safety checks scoring candidate trajectories against drivable area, occupancy, and comfort criteria. We argue that such verifiers are good rewards, and that policies can learn from them at test time. We use verifiers for rejection sampling and optimize driving policies with one gradient step toward accepted trajectories. Our method requires no human demonstrations and no differentiable simulator. It is agnostic to where the samples come from: with a teacher policy it performs online distillation, with the policy's own samples online self-distillation. We show that both regimes support test-time fine-tuning of pre-trained Alpamayo 1.5, AutoVLA, and OneVL policies, improving composite driving scores by up to 24% and lowering displacement error by up to 42%. VADv2, CILRS, and CoverNet policies can even be learned from scratch at test time, given a reasonable initial trajectory distribution. On the KITScenes LongTail and the Waymo Open E2E dataset, our method compensates for domain shifts from nominal to unusual driving scenarios, outperforming all submissions to the 2026 KITScenes Few-Shot Challenge.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.