acceptodds
Under review as a conference paper at ICLR 2027

UniRGBA : Unifying RGB and Alpha in Pretrained Video Diffusion Models

Abstract

Native alpha is essential for compositing visual effects, yet video diffusion models remain RGB-only and matting pipelines underrepresent continuous transparency. We present UniRGBA, an interface-preserving extension of a pretrained video diffusion model for native RGBA video. UniRGBA-VAE expands the Wan2.1 3D causal VAE from RGB to RGBA by zero-initializing the new boundary channels and updating only symmetric 2+2 layers (about 5% of VAE parameters), while retaining the original 16-channel latent interface. A KL-anchored two-stage objective learns alpha structure and sharp compositing details; standard DoRA then adapts the shared Wan2.1 DiT for RGBA T2V and I2V. On a held-out visual-effect benchmark, UniRGBA-VAE improves PSNR from 25.00 to 29.16 over a retrained Wan-Alpha baseline (+4.16dB), including +5.33dB RGB and +3.95dB alpha gains. RGBA T2V improves text alignment (+0.28), naturalness (+0.33), and temporal flickering (+0.0050) over Wan-Alpha, while RGB transfer exposes a measurable -0.095 Dynamic Degree cost. These results show that native alpha can be added through a small VAE-side intervention while retaining a reusable latent interface.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.