acceptodds
Under review as a conference paper at ICLR 2027

Train Once, Rewire Later: Execution-graph Polymorphic Transformers

Abstract

Transformer checkpoints are tied to the *execution graph* used during training. Compilers can optimize how a fixed graph is executed, but changing the graph itself typically requires training a new model. Even a minimal rewrite exposes this dependence: executing attention and the feed-forward layer in parallel raises the language-modeling loss of a 120M transformer from to bits per byte (BPB), despite preserving the same parameters and operators. We call this sensitivity to layerwise execution changes *execution-graph rigidity* and show that it is learned rather than architectural. We introduce *Execution-graph Polymorphic Transformers* (EPTs), which instead train a single set of shared weights across a reference execution graph and alternative layerwise programs, with a consistency objective that aligns their predictions across executions. For the controlled sequential-to-parallel rewrite, the all-parallel program costs at most BPB from 1B to 7B; after 10B training tokens, downstream accuracy differs by at most percentage points across four tasks. We then extend the same objective to six modes per layer: sequential, parallel, reverse, attention only, feed-forward network only, and skip. Measurement-guided search on one 1B checkpoint finds different programs for Blackwell, H100, and L40S, reaching , , and prefill speedup. Programs selected under a BPB search-window budget cost – BPB on disjoint validation contexts. EPTs therefore turn the execution graph from a fixed training choice into a budgeted compiler decision, allowing one training run to support different hardware deployment strategies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.