acceptodds
Under review as a conference paper at ICLR 2027

Visibility Invariance Is Not Enough for RGBA Representation Learning

Abstract

Variational autoencoders (VAEs) for transparent image generation must preserve both appearance and opacity. RGB values at zero opacity do not affect rendering, and premultiplication removes this invisible-color variation exactly. Does this make RGBA autoencoding better? We compare RGBA input representations by adapting two pretrained RGB VAEs, SDXL and Ostris. Three transformations with the same visibility invariance produce different reconstruction quality. After separate validation-based tuning, premultiplication increases alpha mean squared error (MSE) on held-out video frames from the training-source collection by 36.63% for SDXL and 25.68% for Ostris. On human photographs, the same SDXL models instead favor premultiplication, with 70.91% lower alpha MSE and no retraining; the Ostris difference remains uncertain. In SDXL generation with 40 Euler steps and guidance 5, the video-selected premultiplied VAEs improve image-distribution and text-alignment scores despite their worse video opacity reconstruction. These findings neither establish more accurate generated alpha maps nor extend the gains to other sampling settings. Removing invisible RGB guarantees input invariance, but its effect on reconstruction and generation must be evaluated for the intended model, data domain and use. Code is available at https://anonymous.4open.science/r/RGBA-Invariance

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.