Modeling Spatial Transformations Instead of Memorizing Transformed Appearances
Abstract
Learning spatial transformations from transformed images can require surprisingly large amounts of data. We study this effect using controlled affine transformations of MNIST and examine how it changes with dataset size, model capacity, and generalization to unseen appearance–transformation combinations. We evaluate several classifier families, including Vision Transformers, ConvNeXts, Spatial Transformer Networks, and group-equivariant CNNs, spanning substantially different spatial inductive biases. We propose a Transform-Parameterized Implicit Neural Representation (TP-INR), which learns appearance while explicitly representing affine transformation structure and inferring its parameters through reconstruction-error minimization. Across architectures, TP-INR preprocessing substantially improves sample efficiency, with the largest gains in low-data regimes. We further find that a ViT can fail to generalize a known transformation distribution to known classes when their combination was not observed during training, while TP-INR partially mitigates this failure. These results suggest that explicitly encoding known transformation structure can make recognition more data- and parameter-efficient.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.