acceptodds
Under review as a conference paper at ICLR 2027

Raw Image Patches in the Residual Stream: A Data-Efficient Isotropic Vision Transformer

Abstract

In cases where data or compute is constrained, state-of-the-art computer vision models have tended to be anisotropic, i.e. architecturally constrained to regard a different granularity of image detail in different layers. These are Convolutional Neural Networks (CNNs) or Transformer-based models incorporating CNN stems or inductive biases inspired by CNNs. This contrasts with the original ViT, which is isotropic in the sense that every layer represents the image as the same number of constant-size patches. Here we present an ablation study leading to a modernised, post-norm ViT that is isotropic and yet competitive with anisotropic Transformers in data- and compute-constrained settings. In particular, we match or exceed previous Transformer-based models when training for 300 epochs on CIFAR-100 and are competitive with similarly-sized SOTA models when training for 300 epochs on ImageNet-1k. The strength of our model comes partly from incorporating ideas proven effective since the original ViT paper 2020_dosovitskiy_vit and partly from three novel architectural choices. First, our model uses a novel re-weighting of the add-and-norm step of Transformers, allowing stable training across a wide range of depths. Second, we use a Transformer feed-forward block as a stem in place of the usual patch embedding and, third, we pass RMSNormed image patches directly into the residual stream.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.