acceptodds
Under review as a conference paper at ICLR 2027

AlphaRAE: Joint Color and Opacity Modeling with Representation Autoencoders

Abstract

Transparent images require a representation that preserves foreground color and opacity even when a pretrained visual encoder observes only RGB composites. We introduce AlphaRAE, which augments frozen visual features with a spatial color–opacity code and decodes both through a shared RGBA interface. A conditional adapter supplies this interface from a frozen RGB diffusion transformer. At 1024 pixels, removing spatial detail costs 2.91 dB PSNR, while an alpha-only detail input worsens rendered and direct-opacity fidelity at the reported endpoint. On 434 locally reserved images, the high-resolution model reaches 38.83 dB PSNR and 0.01463 LPIPS, compared with 38.52 dB and 0.02123 for AlphaVAE. The 256-pixel comparison remains mixed, although both configurations reduce direct alpha error on local and external images. At matched latent and trainable budgets, the joint model exceeds an independent-alpha control by 0.60 ± 0.12 dB PSNR across three seeds. Same-RGB generation and recovery studies show that the shared decoder can also receive predicted joint latents; a 53-rater study favors its output over same-alpha color refinement on 64 fixed conditions. These results separate codec fidelity from the ability of a generator to supply its latent.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.