Enhancing In-context Panoramic Generation via Geometric-aware Pretraining
Abstract
Panoramic image generation must preserve scene structure across a full spherical field of view, but horizontal wrap-around and latitude-dependent distortion make scene completion and editing difficult. We present Canvas360, a two-stage framework that learns a geometry-aware panoramic prior with parallel RGB–depth pretraining and transfers it to a unified RGB-only generator for four in-context tasks, namely style transfer, inpainting, outpainting, and editing. The pretraining stage combines pseudo-depth supervision, modality-specific representations, and panorama-aware boundary handling to capture scene structure while respecting equirectangular adjacency. To support this two-stage design, we construct Canvas360Dataset with annotated RGB–depth panoramas and task-specific context–target pairs for prior learning and downstream adaptation. On text-to-panorama generation, Canvas360 reduces panorama-specific distributional distance by 49.4% relative to the strongest baseline and outperforms the evaluated baselines in semantic alignment, perceptual quality, and boundary consistency. Completion and editing evaluations further show improvements in perceptual similarity, panoramic fidelity, and reconstruction accuracy. These results indicate that geometry-aware pretraining provides a transferable prior for panoramic generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.