UniRGBA : Unifying RGB and Alpha in Pretrained Video Diffusion Models
Abstract
Native alpha is essential for compositing visual effects, yet video diffusion models remain RGB-only and matting pipelines underrepresent continuous transparency. We present UniRGBA, an interface-preserving extension of a pretrained video diffusion model for native RGBA video. UniRGBA-VAE expands the Wan2.1 3D causal VAE from RGB to RGBA by zero-initializing the new boundary channels and updating only symmetric 2+2 layers (about 5% of VAE parameters), while retaining the original 16-channel latent interface. A KL-anchored two-stage objective learns alpha structure and sharp compositing details; standard DoRA then adapts the shared Wan2.1 DiT for RGBA T2V and I2V. On a held-out visual-effect benchmark, UniRGBA-VAE improves PSNR from 25.00 to 29.16 over a retrained Wan-Alpha baseline (+4.16dB), including +5.33dB RGB and +3.95dB alpha gains. RGBA T2V improves text alignment (+0.28), naturalness (+0.33), and temporal flickering (+0.0050) over Wan-Alpha, while RGB transfer exposes a measurable -0.095 Dynamic Degree cost. These results show that native alpha can be added through a small VAE-side intervention while retaining a reusable latent interface.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.