acceptodds
Under review as a conference paper at ICLR 2027

Batch Multiplexing Language Models

Abstract

Autoregressive decoding repeatedly applies the same model to many independent sequences. We study batch multiplexing, a small change to a decoder-only trans- former that lets several sequences share part of an attention computation while still producing one next-token distribution per sequence. In selected middle layers, a learned multiplexer combines Ksame-position hidden states into one attention in- put. An ordinary attention block runs once per group, and a demultiplexer returns its update to the original sequences before each sequence runs its own MLP. We compare a family of depth-scaled batch-multiplexed models against nanochat us- ing strictly controlled protocols. Across model depths and training horizons, batch multiplexing improves the validation-quality/decode-throughput frontier, which suggests that it is a promising direction for further exploration.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.