Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
Abstract
As large language models serve ever more requests, cumulative inference cost is growing relative to the one-time cost of training. In common serving regimes the two inference phases place different demands on hardware: prompt prefill runs in parallel and tends to be compute-bound, whereas autoregressive decode is sequential and memory-traffic-bound. Conventional width or depth scaling raises both costs together, since every added layer is evaluated in both phases and enlarges the weights read at every decode step. We instead ask whether additional learned computation can be allocated to continuation prediction while preserving the prompt-wide primary computation and a single persistent key–value (KV) cache. We realize this separation with the Decode-Branch Transformer. Its primary path is a complete causal language model that alone processes the prompt and writes the KV cache. At inference, the decode branch is omitted at earlier prompt positions and activated at the final prompt position and each generated position, where it adds continuation-prediction computation without writing persistent state or influencing the primary path. The paths share all major attention, MLP, and output matrices and use separate token embeddings with lightweight coupling. Grouped decode reuses each loaded weight tile and primary-cache region across both paths, so the added arithmetic does not proportionally increase the dominant memory traffic or decode latency. Across matched-training-token comparisons, Decode-Branch attains lower validation loss across architectures and data configurations. A training-compute control further favors interacting trajectories over additional tokens. In MoE models, its structural separation makes the primary and branch expert fan-outs independent knobs for trading prompt cost, continuation cost, and predictive quality. We study two allocation regimes: fixing prefill expert computation while increasing decode computation, and fixing decode expert computation while reallocating the expert budget between the two paths. Together, these experiments expose a prefill–decode–quality trade-off and establish a structural opportunity for phase-specific expert allocation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.