acceptodds
Under review as a conference paper at ICLR 2027

Output-Head Factorization Improves Language Models at No Inference Cost

Abstract

Existing interventions on a language model's output head, including weight tying, learned output maps, mixture models, and low-rank variants, often combine changes to parameter sharing, rank, initialization, or function class. We separate parameter sharing from training-time factorization with matched interventions that change one named axis at a time. Multiplicative factorization of the output head substantially lowers perplexity across the evaluated corpora and architectures, yet folds to a single output projection before inference with no factorization-specific computation. Across dense tied and untied references, perplexity differs by at most 1.8%, whereas the best factorized endpoint is 14.9% lower than its initialization-matched direct reference. The main causal result holds the initial output matrix fixed and changes only its factor split, lowering perplexity by 7.5%. Direct addition provides strong alternatives but does not reach the best factorized endpoint. A local update diagnostic characterizes how factorization changes the output-head update; the endpoint effect is established by intervention. Together, these results show that training-time parameterization of the output head is a consequential deployment-neutral design choice.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.