acceptodds
Under review as a conference paper at ICLR 2027

Progressive Branch-Merge:Trainable Sequence Processing Across Frozen Transformer Depth

Abstract

Equally sized additions to a frozen language model can learn different computations: some widen updates to existing states, while others transform token context throughout decoder depth.We introduce Progressive Branch-Merge (PBM), which pairs each frozen decoder block with a trainable branch that attends to earlier tokens and writes the fused backbone–branch output into a shared residual state that both paths read at the next stage.After one pass through the training text in mathematics, physics, chemistry, and biology, PBM has the lowest mean next-token loss among the tested adaptations of Pythia-1B.On separately fixed fresh evaluation arrays after two passes, it also lowers next-token negative log likelihood relative to a nearly parameter-matched branch-free updater in mathematics, physics, and biology, with four-domain means of 2.0781 and 2.0847; the direct updater has lower loss in chemistry.PBM's relative gain rises with sequence difficulty estimated from thirteen other adaptations, and bypassing its trained branch blocks raises loss more on harder sequences.These results establish stagewise trainable causal processing as a useful allocation of adaptation parameters and show where that allocation yields the larger loss reduction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.