acceptodds
Under review as a conference paper at ICLR 2027

Learning Manifold-Faithful Representations

Abstract

Autoencoders learn latent representations for reconstructing observed data. However, reconstruction training constrains the decoder only at encoded data points, providing no guarantee that perturbations of these codes will decode to samples that remain on the data manifold. Consequently, even a small step away from an encoded point can push the model into an unconstrained region of the latent space, causing the decoded output to degrade. To address this limitation, we introduce manifold-faithful traversal, which constrains movement through the latent space to follow the learned data manifold. By keeping the trajectory on the manifold, latent codes can be repeatedly perturbed while producing plausible intermediate samples without accumulating off-manifold drift over successive steps. We propose a two-phase training framework that first learns a reconstructive representation and then trains the decode–encode mapping to correct perturbed latent codes toward the manifold. Because the true latent manifold is unknown, we approximate this projection using the nearest clean latent code in each training batch. Unlike disentanglement, which seeks latent directions associated with identifiable semantic factors, manifold-faithful traversal does not require specifying what changes along a trajectory; it only requires that the trajectory remain on the data manifold. Experiments on synthetic data and Colored MNIST show that the proposed technique substantially reduces accumulated off-manifold drift and produces coherent trajectories in both random and target-directed walks. We further show that a learned traversal policy together with a memory-efficient training scheme lets the approach scale to larger latent widths and longer trajectories, remaining feasible and effective where the original mechanism is not. Such stable multi-step traversal potentially enables latent space exploration for generative modeling and downstream tasks such as data augmentation that depend on outputs remaining realistic.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.