Panorama-Aware Diffusion Transformer for 360° Video Frame Interpolation
Abstract
Panoramic Video Frame Interpolation (VFI) poses unique challenges due to the equirectangular projection (ERP), which leads to latitude-dependent spatial distortion and a discontinuous longitude boundary. Furthermore, the presence of large and nonlinear motion in ERP panoramas complicates correspondence estimation and temporally coherent synthesis. In this work, we present PADiT, the first Panorama-Aware Diffusion Transformer-based framework specifically designed for video frame interpolation. PADiT introduces new DiT blocks that blend Motion perception and Temporal-attention (MT-DiT), where each transformer block employs multi-head temporal attention to aggregate information from the input ERP frame pair and the latent representation of the intermediate frame. We also integrate flow-derived motion embeddings with diffusion timesteps and frame differences, allowing for the modulation of the denoising blocks through adaptive layer normalization. To enhance geometric and motion modeling for ERP panoramas, we implement spherical position encoding, flow-aware attention, and panorama-aware pyramid feature fusion. These components help mitigate issues associated with ERP, such as positional bias, unreliable motion correspondences, and the misalignment of cross-scale features.Experiments on several widely used panoramic benchmarks, including WEB360, 360VDS, and FlowScape, demonstrate that the proposed PADiT outperforms state-of-the-art methods in panoramic VFI. Specifically, it reduces the FloLPIPS and LPIPS metrics by up to 9.33% and 8.44%, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.