Controlling Information Flow in Vision Transformers with Variational Bottlenecks
Abstract
We make the information transmitted by attention in vision transformers a measurable and controllable quantity. By inserting variational information bottlenecks on attention-mediated writes to the residual stream, we train models under an explicit information cost, producing a spectrum from independent patch processing to fully expressive global attention. The same mechanism is compatible with both supervised classification and DINO-style self-supervised learning, allowing us to study how useful visual representations emerge as progressively more information is exchanged across patches. On ImageNet-100, we characterize how performance and representation quality vary with information rate, and how communication is allocated across depth, attention heads, and spatial locations. Under tight information constraints, only a small number of attention heads become active, exposing simple visual computations that precede richer global representations. These results provide a direct way to control internal information flow and yield vision transformers whose computations are more tractable for mechanistic analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.