acceptodds
Under review as a conference paper at ICLR 2027

Calibration-Guided One-Shot Vocabulary Pruning For Edge LLMs

Abstract

Modern language models often employ a large vocabulary to cover multiple languages and modalities, which makes their deployment expensive. To address this challenge, we study static vocabulary-head pruning that permanently restricts the head to a survival subset chosen once, offline, with no post-pruning fine-tuning. The de-facto criterion for this choice is token frequency, and we first establish why frequency is the wrong signal along three axes: a participation-frequency mismatch, blindness to substitutability, and the loss of rare-but-critical tokens. We introduce (Participation And Substitution Selection), a calibration-guided, training-free method built on two complementary per-token signals: participation which measures the rate at which a token enters the model's top- at decode time; and substitution cost} which is a mass-weighted logit deficit incurred by replacing a token with its nearest surviving neighbour. We then propose a two-stage selector that first uses participation to tag tokens as keep, prune or uncertain decisions, and then uses substitution cost to rank only the remaining uncertain tokens. reduces the decode-time head cost from to per token and at a keep fraction, retains nearly entire of full-head quality on multiple evaluation benchmarks as GSM8K, TrivaQA, MATH-500, etc.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.