HorizonLoRA: Distilling Motion Adapter Banks for Trajectory-Conditioned Video World Models
Abstract
Long-horizon video world models must accommodate changing camera trajectories while reusing a large frozen visual prior. Static LoRA applies one update throughout a rollout, whereas motion-specialized adapter banks retain multiple sets of expert weights. We propose HorizonLoRA, which distills a motion adapter bank into a shared dictionary of normalized rank-one updates and predicts bounded, chunk-specific coefficients from the initial frame, text prompt, ordered camera controls, cumulative commanded pose, and rollout time. A two-stage procedure first fits the dictionary to individual teachers and their mixtures, then trains the coefficient generator with video supervision and Gram-space parameter distillation. The teacher bank is discarded at deployment. We compare five adaptation strategies under 1-chunk, 4-chunk, and 16-chunk recursive rollouts. In the reported comparison, HorizonLoRA reduces deployed adapter storage by 72.4% relative to the learned bank and records lower 16-chunk similarity-aligned trajectory error on the evaluated Sekai and DL3DV samples. Other adapters obtain stronger appearance or boundary-continuity scores. These results characterize the geometry–appearance–storage trade-off of teacher-distilled conditional adaptation for recursive video generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.