Inversely Autoencoding Multimodal Data into a Shared Latent Space
Abstract
Learning multimodal representations in a shared latent space is a fundamental challenge towards unified models. In this work, we introduce the Inverse Autoencoder (Inv-AE), a novel framework based on the insight that decoding a shared latent representation into disparate modalities is far more tractable than encoding multimodal data into an aligned space. Inv-AE first optimizes shared latent representations alongside multiple modality-specific decoders, followed by training encoders to map heterogeneous inputs into this established feature space. To accommodate the disparate information densities across modalities, we propose the Latent Representation Sequence (LRS), a multi-scale hierarchical sequence of tensors with progressively increasing sizes, allowing each decoder to selectively leverage an adaptive prefix length. Consequently, low-information-density to high-information-density autoencoding (e.g., images to labels) naturally maps as a classification task, whereas high information density to low-density autoencoding utilizes a Transformer to compensate for finer details via conditional autoregression. Experiments on standard benchmarks demonstrate that our approach delivers competitive performance across multiple tasks while remaining highly extensible to expert architectures with minimal modification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.