Self-Reconstruction Dynamics for Autoencoder Reconstruction Refinement
Abstract
Standard autoencoder (AE) inference uses a single encoder–decoder pass, although the resulting latent representation need not be optimal for each sample under the fixed decoder. This raises a natural question: Can a trained AE itself reveal information useful for improving its own reconstruction? We investigate this question by repeatedly applying a frozen AE to its own reconstruction, producing transient image- and latent-space trajectories that we call Self-Reconstruction Dynamics (SRD). Although repeated self-reconstruction degrades fidelity in the AEs studied here, the resulting SRD contains useful sample-specific information for correcting the initial reconstruction. We therefore propose SRD-guided Reconstruction Refinement (SRD-RR), which predicts a latent correction from a short SRD while keeping the AE frozen and requiring no per-sample test-time optimization. We also introduce an MSE recovery ratio (MSE-recov) relative to an empirical decoder-optimized reference. Across six image datasets, SRD-RR recovers an average of 38.6% of the empirically recoverable MSE gap with one trajectory transition and 45.3% with two. A two-transition variant trained without direct access to the original images, using an SRD-derived pseudo-target, achieves 40.7% recovery and a 1.74dB average PSNR gain. Removing trajectory information substantially reduces the gain, while cross-sample trajectory assignment causes severe degradation, showing that the useful SRD information is strongly sample-specific. Nonlinear SRD-conditioned refinement also consistently outperforms both fixed and trained linear latent correction. We further evaluate SRD on a pretrained DINOv2-based representation autoencoder (RAE) with substantially different latent dynamics. SRD conditioning again improves a matched trajectory-free predictor, showing that trajectory information remains useful in this different AE setting. However, pixel-MSE latent refinement exposes a strong mismatch between pixel fidelity and perceptual quality, while the SRD-derived pseudo-target substantially mitigates this perceptual degradation. Together, these results establish SRD as a useful sample-specific signal for reconstruction refinement, while showing that the choice of refinement objective determines how this information translates into pixel and perceptual quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.