Rare-Aware Autoencoding: Reconstructing Spatially Imbalanced Data
Abstract
Autoencoders are widely used for unsupervised and self-supervised learning when labels are scarce, but their reconstruction objectives can be biased by spatially imbalanced image data, favoring dominant background or frequent pixel patterns. This issue is common in scientific imaging including medical imaging, biology, physics, and astronomy, where relevant image content may appear sparsely or non-uniformly across spatial locations, and manual annotations are limited or expensive. We propose a reconstruction framework with two components: a self-information-based loss that upweights statistically uncommon pixel intensities conditioned on fixed spatial locations, and Sample Propagation, a replay strategy that re-exposes the model to hard-to-reconstruct samples across batches. Unlike most data-imbalance methods, which rely on task-specific annotations, our approach operates in the unsupervised setting. Nevertheless, we also adapt and benchmark methods from supervised learning. Experiments on a controlled simulated dataset and three real-world datasets spanning physical, biological, and astronomical domains show improved reconstruction over existing baselines, particularly in settings with strong spatial imbalance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.