LW-MIH: GENERATIVE MULTI-IMAGE HIDING WITH VAE PACKING AND LEARNABLE WHITENING
Abstract
Multi-image hiding conceals several secret images in one image sent over a public channel, but existing methods face a critical issue: cover-based methods reach high capacity yet modify the cover and leave traces, whereas coverless generative methods are secure yet hide only a few images, and neither states how many images one cover can carry nor how to trade that number against per-image fidelity. To address this issue, we propose LW-MIH, a generative coverless multi-image hiding framework with high capacity and high recovery fidelity. LW-MIH first encodes every secret with a frozen foundation VAE and packs the latents into a fixed-size tensor, so the number of images one cover carries follows from the encoder geometry: up to 96 with Sana, 48 with SD1.5 and 12 with FLUX. It then fills the tensor in one of three ways — full capacity, high-frequency dual-slot and upsampling — turning the number of hidden images and the fidelity of each into a discrete set (FLUX 12/6/3, SD1.5 48/24/12/3, Sana 96/48/24/6). Finally, an invertible neural network maps the packed tensor to the noise end of a frozen rectified flow, and a learnable invertible whitener aligns the field with the flow's Gaussian prior, adding no capacity cost and keeping recovery exact. Experimental results show that at full capacity the 16-bit PSNR of the recovered secrets is 35.03, 28.17 and 26.89 dB for FLUX, SD1.5 and Sana, within 0.15, 0.07 and 0.06 dB of their VAE ceilings, and rises to 41.03, 37.08 and 37.91 dB when fewer images are hidden; the 8-bit values are 24.12, 22.77 and 20.18 dB. The whitener gains 3.97 and 3.36 dB of 8-bit PSNR over a random orthogonal mixer of the same form on FLUX and SD1.5, and at full capacity leaves a worst gap of 0.41 dB where the pipeline without a mixer falls 12.4 dB below its encoder's limit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.