Local Splits, Global Shifts: An Interaction Decomposition of Retokenization in Language Models
Abstract
Different tokenizations of identical bytes can change a language model's predictions. We introduce the first exact local-to-global framework for attributing same-byte retokenization sensitivity to contiguous spans and their interactions. Universal cuts split the tokenization space into the finest independent contiguous blocks. Summing next-token probabilities by first byte makes predictions comparable across tokenizations. Functional analysis of variance separates first-order effects of single blocks from multi-block interactions. Resampling a single block measures its local influence, whose sum upper bounds global variation. Their ratio, the interaction order, is one when effects add and exceeds one when blocks interact. An exhaustive binary-block analysis finds mostly first-order effects with substantial interactions. Across prose, code, and multilingual text, local influence predicts global variation beyond severity and other controls, and prompt-pooled interaction order rises with severity. Independent interventions show that local influence ranks fragile spans whose retokenization is more likely to induce task failures. We derive Local Retokenization Consistency (LRC), a symmetric one-block objective that penalizes local influence while preserving the expected loss of stochastic supervised fine-tuning. In matched-step adaptation of an instruction-tuned model, LRC achieves the highest mean canonical and retokenized accuracy, averaged equally over tasks, and the smallest mean gap among the compared objectives.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.