acceptodds
Under review as a conference paper at ICLR 2027

Learning to Retokenize: Improving Downstream Performance of Frozen Language Models

Abstract

Language modelling in LLMs begins with tokenization: splitting the input string into tokens drawn from a fixed vocabulary according to predefined rules. Because the tokenizer is trained on a generic corpus before model pre-training and is not tailored to any downstream task, its tokenization strategy can be ill-suited to certain domains or tasks, hindering performance. Adapting tokenization per task could help, but retraining the tokenizer on a new corpus likely changes the vocabulary and requires model adaptation, which is costly in general and infeasible for closed-weight models. A tokenizer that reuses the original vocabulary, however, could be applied to the model seamlessly. In this work, we explore training a task-specific tokenizer that is constrained to the original vocabulary of a given pre-trained LLM. To this end, we propose the retokenizer, a sequence-to-sequence model that maps an input token sequence to a new one within the original vocabulary, while preserving the underlying string content. During training, our retokenizer learns from a task-specific reward and only requires black-box access to the language model, making the approach suitable for closed-weight models. Across arithmetic, character manipulation, and code understanding tasks, our learned retokenization improves accuracy over the default tokenization by 10% on average. Moreover, we show that in some settings the retokenizer uncovers intuitive task-dependent tokenization rules, such as right-to-left digit tokenization for arithmetic. These learned rules suggest that, given the right objective, effective tokenization strategies can be discovered automatically rather than designed by hand.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.