Train Once, Rewire Later: Execution-graph Polymorphic Transformers
Abstract
Transformer checkpoints are tied to the *execution graph* used during training. Compilers can optimize how a fixed graph is executed, but changing the graph itself typically requires training a new model. Even a minimal rewrite exposes this dependence: executing attention and the feed-forward layer in parallel raises the language-modeling loss of a 120M transformer from to bits per byte (BPB), despite preserving the same parameters and operators. We call this sensitivity to layerwise execution changes *execution-graph rigidity* and show that it is learned rather than architectural. We introduce *Execution-graph Polymorphic Transformers* (EPTs), which instead train a single set of shared weights across a reference execution graph and alternative layerwise programs, with a consistency objective that aligns their predictions across executions. For the controlled sequential-to-parallel rewrite, the all-parallel program costs at most BPB from 1B to 7B; after 10B training tokens, downstream accuracy differs by at most percentage points across four tasks. We then extend the same objective to six modes per layer: sequential, parallel, reverse, attention only, feed-forward network only, and skip. Measurement-guided search on one 1B checkpoint finds different programs for Blackwell, H100, and L40S, reaching , , and prefill speedup. Programs selected under a BPB search-window budget cost – BPB on disjoint validation contexts. EPTs therefore turn the execution graph from a fixed training choice into a budgeted compiler decision, allowing one training run to support different hardware deployment strategies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.