RolloutRemover: Scribble-Guided Object Removal via Latent Rollout Modeling
Abstract
Object removal aims to remove target objects and recover the background they occlude. Existing methods have achieved strong visual fidelity, but they still rely on mask-conditioned inpainting and therefore require an accurate object mask as input. In practice, however, users often provide only a casual scribble, while current pipelines either rely on limited scribble generalization or introduce external models to bridge the gap. We present **RolloutRemover**, a video diffusion framework for scribble-guided object removal. Instead of directly predicting the final removed image, RolloutRemover interprets the scribble as removal intent and performs removal through a structured latent rollout that explicitly disentangles object corruption from background restoration, providing stronger supervision over the removal process. We also introduce **SGORBench** with multiple mask sets for comprehensive evaluation of image object removal. Experiments show that RolloutRemover achieves strong removal quality and is substantially more robust to sparse user-provided scribbles than prior mask-dependent pipelines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.