HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation
Abstract
Closed-loop driving simulation requires real-time interaction beyond short offline clips, pushing current driving world models toward autoregressive (AR) rollout. As generated frames become the context for subsequent chunks, prediction errors accumulate over time. The key difficulty is that a teacher trained only on clean histories provides unreliable corrective guidance when conditioned on the student's generated frames. A natural question is: can a teacher provide reliable corrective guidance throughout a student's long AR rollout? We address this with HorizonDrive, a teacher-centric training-and-distillation framework for AR driving simulation. Through scheduled rollout recovery (SRR), we mine semantically degraded yet layout-preserving states from long driving rollouts, turning their retained scene structure into corrective supervision for training a recovery-capable teacher. Further, teacher rollout distillation (TRD) transfers this capability to a short-chunk, few-step student through fixed-window distribution matching, enabling continual correction throughout long AR rollouts with bounded memory and per-window computation. HorizonDrive natively supports minute-scale AR rollout under bounded memory; on nuScenes, HorizonDrive reduces FID by 52% and FVD by 37%, and lowers ARE and DTW by 21% and 9% relative to the strongest long-horizon streaming baselines. Open- and closed-loop evaluations further show improved controllability and downstream planning utility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.