What is this Object and Where is it Going?
Abstract
Video outpainting aims to extend content beyond the original frame boundaries while preserving inter-frame and intra-frame consistency. Existing methods have achieved promising results, but still fall short in complex scenes and motions. In particular, video outpainting inherently involves partially observed objects, posing significant challenges in maintaining object-level consistency in both appearance and motion. To address these issues, we propose RAVO, a retrieval-augmented test-time adaptation framework that complements incomplete object observations with relevant videos to support object-consistent video outpainting. However, di- rectly exploiting retrieved videos is non-trivial due to discrepancies in visual char- acteristics and dynamics between the retrieved and source objects. To bridge this discrepancy, we introduce Object-Specific Adaptation (OSA), which learns trans- ferable part-level correspondences between source and retrieved objects while retaining source-specific identity cues. Furthermore, conventional 3D attention tightly couples spatial and temporal interactions, limiting the effective incorpora- tion of heterogeneous object-level cues during adaptation. We therefore propose Spatial-Temporal Decoupled Adaptation (STDA), which separately adapts spatial appearance and temporal dynamics to mitigate their mutual interference. Beyond these methodological contributions, we introduce ObjectBench to address the lack of benchmarks for object-level evaluation, providing an object-centric assessment of appearance and motion consistency. Extensive experiments show that RAVO achieves state-of-the-art performance on standard benchmarks and consistently outperforms existing methods on ObjectBench, demonstrating broad efficacy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.