Aligning and Folding: What Networks Do Before They Learn
Abstract
Architectures are usually organized by name, which hides an elementary question: what are the reusable pieces from which they are assembled? We answer that a network is an alternation of two processes. An align exposes patterns on arbitrary supports in comparable coordinates; a fold combines responses and may identify distinctions. Convolutional networks, transformers, and state-space mixers are sub-families of this single class: residual blocks, a pre-norm transformer block, and the mixer of an official Mamba-130M agree with their translated programs to numerical precision on the order of in forward values and per-parameter gradients, and covering the state-space mechanism costs one additional align primitive. The class is built to answer two questions before training: whether a given architecture can realize a task within a declared budget, and which architecture in a predeclared family should be chosen. Two empirical principles make this precise. Information is pattern: a task is generated by finitely many admissible patterns of a label-independent grammar. Failure is cost: once an architecture can jointly expose those patterns, an exact realization of the task is reachable from some admissible initialization whenever the declared budgets suffice. Each has a stated falsifier. For diagnosis, a representation admits an unrestricted deterministic readout precisely when its fibers refine the label partition, and cross-label collisions are irreversible downstream. We define a capacity that counts, as a formal upper bound, the ways an architecture can transform the basis of the patterns it jointly exposes. Predictions made before training hold on controlled tasks, including a certified impossibility and a targeted repair. For selection, the same capacity and cost are computed for every member of a predeclared family without training candidates, returning a finite cost-Pareto set; the framework finds a vision transformer with 25% fewer parameters than ViT-Lite-7/4 that matches ViT-Lite-6/4, and, among inherited-weight subnets of AutoFormer's ImageNet search space, one with fewer parameters matching the reported accuracy of AutoFormer-T.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.