BRAVER: Animating RGB References into Reusable RGBA Videos
Abstract
Creators often begin with an image or video that captures the appearance or motion they want to preserve. Turning these references into reusable animated assets requires retaining the intended foreground while separating it from its original background, including through fine boundaries and translucent regions. This calls for both a faithful transparency representation and reference conditioning that distinguishes the subject from its surroundings. We introduce BRAVER (Background-Robust Animation of Visual Elements for Reuse), a matte-free framework for text-conditioned generation, image-guided animation, and aligned foreground extraction from RGB videos. BRAVER combines a pretrained video prior with a compact RGBA autoencoder. Regional alpha supervision complements reconstruction over opaque interiors, boundaries, and translucent regions. Two-stage adaptation establishes an RGBA generation prior and incorporates visual inputs through shared adaptation and task modulation. Training references are composited over varied backgrounds while their original RGBA assets remain the targets, providing supervision for separating the foreground from its surroundings. The framework requires neither user-provided mattes nor a separate output-matting stage. We also curate AlphaElements, an RGBA video collection spanning animated stickers, motion graphics, and visual effects, to support reconstruction and conditional generation. BRAVER outperforms the evaluated baselines in aesthetic quality and naturalness for text-conditioned generation, and reference consistency for image-guided animation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.