PAATRA: Parameter-Allocation-Aware Training for Small Language Models
Abstract
Small language models (SLMs) often inherit tokenizers designed for much larger model families. In our 64M-parameter GPT-2-style student using Qwen2.5's 152K-token vocabulary, 81.8% of the parameters are token embeddings. We study this vocabulary–transformer allocation with PAATRA, a pipeline that retains a frequency-ranked subset of the teacher vocabulary, distills from the teacher over that subset, and reinvests the freed parameters into transformer width, subject to an approximately fixed total budget. Distilling from Qwen2.5-0.5B on WikiText-103 with a shared training recipe, students with 47K and 20K tokens reach mean bits per byte (BPB) of and over three seeds, 4.9% below the full-vocabulary baseline (). A 10K vocabulary is slightly worse than the baseline (), and a 20K vocabulary without reinvestment is worse still (1.4322). Two results qualify this gain. First, in a single-seed learning-rate sweep, the full-vocabulary baseline tuned to reaches 1.3455, matching the best 20K student. Second, small, genuine BPE/Unigram tokenizers trained with plain cross-entropy under a different protocol achieve lower BPB (1.19–1.27) than all distilled students. We therefore present PAATRA as a controlled measure of the allocation effect under a fixed recipe, for a single teacher and a single corpus. We do not claim that reduced vocabularies win under per-configuration tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.