Position: Transformer Depth Graphs Should Be First-Class Training Variables
Abstract
Transformer depth can be changed through skipping, looping, reordering, and parallel execution, but studies differ in which graph they train, which pretrained weights they edit, and what resources they allow. An operation name therefore does not identify a comparable experiment. We argue that the complete training graph, including its block calls, dependencies, parameter-sharing identities, merge rules, execution order, path-selection rule, and settings attached to individual block calls, should be treated as a shared experimental object. Claims about this object require graph-aware comparisons that separate whether training on the graph learns the intended computation, what quality its cost buys, how fixed weights respond to an execution edit, and what a multi-operation graph adds over simpler alternatives. In four shared runs of a controlled strict-order task, Serial has lower test error than Parallel-concat in all four but lower error than Looping in only one; the Looping comparison is not matched on estimated arithmetic. In language modeling, a wider shared-looping model has lower validation loss than a narrower serial model at similar parameter counts and matched training tokens and block calls, but width and sharing remain confounded and the looping design takes about 1.56 times as long to train. In the tested checkpoints, fixed-weight loop-count edits increase loss, while training with skipping lowers later skipping damage but misses the full-execution quality criterion. These experiments address task fit, quality and cost, and execution dependence, but none trains a multi-operation graph, so its value beyond simpler graphs remains untested. Together, these cases support a comparison contract that can select a proposed graph or retain a serial or simpler alternative.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.