acceptodds
Under review as a conference paper at ICLR 2027

Panorama-Aware Diffusion Transformer for 360° Video Frame Interpolation

Abstract

Panoramic Video Frame Interpolation (VFI) poses unique challenges due to the equirectangular projection (ERP), which leads to latitude-dependent spatial distortion and a discontinuous longitude boundary. Furthermore, the presence of large and nonlinear motion in ERP panoramas complicates correspondence estimation and temporally coherent synthesis. In this work, we present PADiT, the first Panorama-Aware Diffusion Transformer-based framework specifically designed for video frame interpolation. PADiT introduces new DiT blocks that blend Motion perception and Temporal-attention (MT-DiT), where each transformer block employs multi-head temporal attention to aggregate information from the input ERP frame pair and the latent representation of the intermediate frame. We also integrate flow-derived motion embeddings with diffusion timesteps and frame differences, allowing for the modulation of the denoising blocks through adaptive layer normalization. To enhance geometric and motion modeling for ERP panoramas, we implement spherical position encoding, flow-aware attention, and panorama-aware pyramid feature fusion. These components help mitigate issues associated with ERP, such as positional bias, unreliable motion correspondences, and the misalignment of cross-scale features.Experiments on several widely used panoramic benchmarks, including WEB360, 360VDS, and FlowScape, demonstrate that the proposed PADiT outperforms state-of-the-art methods in panoramic VFI. Specifically, it reduces the FloLPIPS and LPIPS metrics by up to 9.33% and 8.44%, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.