acceptodds
Under review as a conference paper at ICLR 2027

Modeling Spatial Transformations Instead of Memorizing Transformed Appearances

Abstract

Learning spatial transformations from transformed images can require surprisingly large amounts of data. We study this effect using controlled affine transformations of MNIST and examine how it changes with dataset size, model capacity, and generalization to unseen appearance–transformation combinations. We evaluate several classifier families, including Vision Transformers, ConvNeXts, Spatial Transformer Networks, and group-equivariant CNNs, spanning substantially different spatial inductive biases. We propose a Transform-Parameterized Implicit Neural Representation (TP-INR), which learns appearance while explicitly representing affine transformation structure and inferring its parameters through reconstruction-error minimization. Across architectures, TP-INR preprocessing substantially improves sample efficiency, with the largest gains in low-data regimes. We further find that a ViT can fail to generalize a known transformation distribution to known classes when their combination was not observed during training, while TP-INR partially mitigates this failure. These results suggest that explicitly encoding known transformation structure can make recognition more data- and parameter-efficient.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.