Encoding The LLM Vocabulary Bottleneck
Abstract
Large Language Models (LLMs) have achieved strong performance across a wide range of natural language processing tasks, but their increasing scale raises significant challenges in computational and memory efficiency. One of the major contributors to these costs is the decision layer (referred to as the Softmax layer), a fully connected projection followed by a Softmax activation. Its learned parameter count and associated optimizer-state count grow linearly with vocabulary size and can account for a substantial fraction of a smaller model's trainable state. In this work, we propose a framework based on Error Correcting Output Coding (ECOC), in which the standard Softmax layer is recovered as a special case corresponding to the identity coding, while more compact decision-boundary encodings are obtained by replacing this identity structure with alternative ECOC encodings, yielding an ECOC-Softmax decision layer. Experiments spanning three architectures and vocabularies of up to 114K tokens show that, with binary codewords of length 200, ECOC-Softmax keeps top-1 accuracy within around 2 percentage points of the best Softmax baseline while reducing the trainable output-layer parameter count by up to 571 times, total trainable parameters by 18–23%, and output-layer FLOPs per predicted token by 3.8–4.4 times. Extending this, ECOC-Tying reduces total trainable parameters and associated optimizer state by 22–29% relative to tied Softmax.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.