Towards Coherent Scene Reconstruction in Video Object Removal
Abstract
Video object removal requires coherent scene reconstruction after eliminating unwanted objects and their associated effects. Cross-frame observations can reveal occluded content, yet mixed input representations, misleading cues, and uneven reconstruction support limit their effective use. We propose TIE, a framework integrating Type-Isolated Attention and Evidence-Guided Supervision for coherent scene reconstruction. Type-Isolated Attention organizes the original video, mask, and noisy target into distinct token streams and directs their interaction, maintaining target access to conditions while preventing noisy-target feedback into condition representations. Beyond accessibility, Evidence-Guided Supervision uses paired edited targets and frozen cross-frame matching to assess whether observations provide valid support for the scene after removal. The resulting evidence calibrates the allocation of complementary point-wise and region-wise auxiliary supervision, emphasizing precise matching where support is strong and local appearance and variation constraints where it is weaker. This allocation supplements dense reconstruction to account for differences in support across locations. Extensive experiments demonstrate that TIE achieves state-of-the-art overall performance, enabling more complete removal of objects and their associated effects, more faithful reconstruction of occluded content, and improved temporal coherence while preserving the surrounding scene.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.