acceptodds
Under review as a conference paper at ICLR 2027

Model Carving: Optimizing the Topology of Standard and Looped Transformers

Abstract

Transformer architectures are largely fixed: the same block repeats across layers, and the connections between them, for both standard and looped transformers, are prescribed in advance rather than learned. We ask: can a Transformer's topology, learned jointly with its weights, beat a hand-designed one on the quality-efficiency frontier? In this paper, we introduce Carving, a method to learn a model's topology during training. Carving places a gate on each edge of a Transformer's computation graph, trains the gates jointly with the weights, and compiles the remaining graph into a sparser, carved model. On both standard and looped Transformers, carved models beat the larger models they were carved from, reaching lower validation loss with half the inference FLOPs, one-third and one-half the KV cache, and and faster decoding throughput at the 1.3B and 2.7B scales, respectively. Carved models also beat training-FLOP-matched shallower Transformers with the same inference cost in validation quality and throughput. The architectures learned share similar features across model sizes, scales, and seeds, removing 35% of the MLP sublayers and 53% of the attention sublayers on average. These results suggest that standard, homogeneously repeated blocks are not always the most efficient design. Carved models also outperform larger models pruned post-hoc to their same inference cost on the quality-efficiency frontier, showing that the learned topology both differs from and beats standard post-hoc compression.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.