Beyond Dense Context: Efficient Video Inpainting via Position-aware Context Extraction
Abstract
Video inpainting is inherently asymmetric, as only the masked region requires synthesis, while the non-mask region primarily provides context. Existing efficient methods either summarize non-mask content into spatially coarse global features or estimate dense spatial features, creating a trade-off between computational efficiency and fine-grained context preservation. To understand which context is needed, we analyze masked-to-non-mask attention in a dense baseline. We find that masked queries concentrate their attention on a small subset of non-mask tokens, particularly those near the mask and in semantically informative regions. Motivated by this observation, we introduce PACE, a video inpainting framework that extracts compact, position-aware context. PACE assigns sparse spatial anchors to a small set of learnable queries. Through attention that incorporates both content and relative spatiotemporal position, these queries aggregate non-mask information from across the video into compact context features. An adaptive attention fusion mechanism then dynamically balances the extracted context with mask-side information. By compressing dense non-mask tokens into compact context representations, PACE accelerates masked-region denoising while preserving the contextual information needed for high-quality inpainting. Experiments on DAVIS, YouTube-VOS, and AdErase, our new advertising video benchmark for text and logo removal, demonstrate strong inpainting quality. PACE outperforms the efficient baselines EditCtrl and YOSE in quality and achieves an average – speedup over them across different mask ratios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.