FastEraser: Achieving Efficient and Effect-Aware Video Object Removal via Sparsity-Driven Distillation
Abstract
Effect-aware video object removal (VOR) aims to eliminate not only a target object but also its associated environmental effects, such as shadows and reflections. Although recent generation-based VOR methods achieve compelling reconstruction quality, their billion-scale architectures and iterative denoising processes incur substantial computational cost. Existing acceleration strategies either apply mask-guided spatial sparsity, which overlooks effects outside the input mask, or directly reduce the number of denoising steps, which introduces noticeable blur and noise. Few-step distillation offers a promising alternative, yet conventional Distribution Matching Distillation (DMD) relies heavily on the asymmetry induced by classifier-free guidance (CFG), making it poorly suited to VOR models that are inherently CFG-free. We propose FastEraser, an efficient framework for effect-aware VOR built upon Sparsity-Driven Distillation (SDD) and Pixel Reconstruction Loss (PRL). SDD replaces CFG-induced asymmetry with a dense-to-sparse attention discrepancy between the real and fake score models, restoring the driving signal required for stable few-step adaptation without restricting removal to the input mask. PRL further anchors the student prediction to the ground-truth video by decoding its denoised latent through a frozen VAE and applying direct supervision in RGB space, thereby preserving removal capability and high-frequency details. Extensive experiments on four benchmarks demonstrate that FastEraser achieves performance comparable to existing 50-step models using only 4 denoising steps, substantially reducing inference cost while maintaining effect removal quality and reconstruction fidelity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.