acceptodds
Under review as a conference paper at ICLR 2027

STRIDE: Scene Transition via Interpolative Diffusion Estimation

Abstract

Video transition generation, the synthesis of coherent intermediate frames bridging a given start and end frame, is a core problem in generative video modeling. Existing diffusion and flow methods are not only computationally heavy but also rely on large pretrained image/video backbones. We present STRIDE (cene ansition via nterpolative iffusion stimation), an inference-efficient bidirectional latent-diffusion framework whose 3D U-Net denoiser (1.54B parameters) is customized and trained as generative-prior-free, that carries no transition or motion prior from any pretrained generator (a frozen VAE and CLIP serve only as fixed interfaces). STRIDE couples position-aware progressive masking, coupled bidirectional temporal attention with a forward/backward consistency loss, and multi-anchor CLIP conditioning, is trained under a four-phase motion-filtered curriculum with an annealed-fusion inference schedule. Operating in a compact pixel and latent space, it infers to faster than DynamiCrafter, SEINE, TVG, and Wan2.1-VACE-1.3B, at to lower per-clip compute than DynamiCrafter, SEINE, and TVG. Further, it attains the best inter-frame structural consistency (VSI, FSIM, and tMS-SSIM on all three benchmark datasets i.e., Morph-Bench, TC-Bench, and TransitBench and tPSNR on one of the three). Although SVD-KFI (Generative Inbetweening) wins the four perceptual/appearance-similarity metrics (CLIPSim, DreamSIM, LPIPS, EMD) on all three benchmarks, its inference is slower than STRIDE, with Wan and Framer occupying intermediate positions. Because all eight metrics are computed between consecutive frames, they measure inter-frame smoothness rather than motion fidelity. This makes the resulting frontier an operating-point trade-off between inter-frame structural consistency and perceptual/appearance similarity, rather than a quality ranking. In our human study (5 models 48 clips 3 raters ratings), we observe high inter-rater agreement (Krippendorff's –). The Friedman omnibus test reveals significant differences in transition-smoothness ratings across models (). Furthermore, Holm-corrected pairwise Wilcoxon tests confirm that STRIDE is significantly preferred over all four pretrained-prior baselines (DynamiCrafter, SEINE, TVG, Wan2.1-VACE-1.3B). Despite using no pretrained diffusion backbone, STRIDE achieves significantly smoother transitions than pretrained-prior baselines, at the cost of occasional reductions in intermediate-frame visual quality near the transition midpoint.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.