acceptodds
Under review as a conference paper at ICLR 2027

Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation

Abstract

As large language models serve ever more requests, cumulative inference cost is growing relative to the one-time cost of training. In common serving regimes the two inference phases place different demands on hardware: prompt prefill runs in parallel and tends to be compute-bound, whereas autoregressive decode is sequential and memory-traffic-bound. Conventional width or depth scaling raises both costs together, since every added layer is evaluated in both phases and enlarges the weights read at every decode step. We instead ask whether additional learned computation can be allocated to continuation prediction while preserving the prompt-wide primary computation and a single persistent key–value (KV) cache. We realize this separation with the Decode-Branch Transformer. Its primary path is a complete causal language model that alone processes the prompt and writes the KV cache. At inference, the decode branch is omitted at earlier prompt positions and activated at the final prompt position and each generated position, where it adds continuation-prediction computation without writing persistent state or influencing the primary path. The paths share all major attention, MLP, and output matrices and use separate token embeddings with lightweight coupling. Grouped decode reuses each loaded weight tile and primary-cache region across both paths, so the added arithmetic does not proportionally increase the dominant memory traffic or decode latency. Across matched-training-token comparisons, Decode-Branch attains lower validation loss across architectures and data configurations. A training-compute control further favors interacting trajectories over additional tokens. In MoE models, its structural separation makes the primary and branch expert fan-outs independent knobs for trading prompt cost, continuation cost, and predictive quality. We study two allocation regimes: fixing prefill expert computation while increasing decode computation, and fixing decode expert computation while reallocating the expert budget between the two paths. Together, these experiments expose a prefill–decode–quality trade-off and establish a structural opportunity for phase-specific expert allocation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.