acceptodds
Under review as a conference paper at ICLR 2027

The Outsized Vocabulary Tax: Selective-Vocabulary Decoding for Edge LLMs

Abstract

Modern small language models pair a compact backbone with a large vocabulary which incur an output projection cost of per decoded token, reaching up to 28% of total parameters and 20% of decode latency on edge hardware. Recent cluster-routed decoding reduces this to without modifying any weights, by scoring only the tokens in the few of clusters whose centroids best match a given hidden state. Existing methods cluster LM-head rows directly, but we show that tokens close in weight space are often not the tokens that similar hidden states select. We introduce (Calibrated Hidden-state Embedding Selection), a training-free method that represents each token by a count-adaptive shrinkage estimate on the unit sphere. This estimate interpolates between the mean direction of the hidden states that select the token and a prior given by its LM-head row. We further propose , a closed-form anisotropic centroid update that aligns clustering with inference-time inner-product routing. On Qwen2.5 and Gemma2 models, reaches 95% top-1 routing accuracy while scoring only 6-8% of the vocabulary and yields up to 3.2 LM-Head speedup on Apple M4 and Intel Xeon CPUs. Across multiple downstream benchmarks (e.g., GSM8K, MATH, TriviaQA), can closely match the full LM-Head performance while utilizing only 18-20% of total vocabulary.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.