When and Where Can Learned Tokenization Be Refined at Test Time?
Abstract
Learned byte-level tokenizers expose an inference-time degree of freedom that fixed subword tokenizers do not: the model can vary how much of its learned latent resolution is used for an observed prefix. We study two forms of test-time refinement—where to restore boundaries inside a reduced native mesh and when to return a prefix to full native resolution. The experiments reveal a clear granularity asymmetry. Fine-grained boundary placement is unstable: dense insertion hurts and first-layer geometry or attention does not yield a reliable long-horizon location rule. Prefix-level allocation is substantially more predictable. A lightweight router that estimates continuation-likelihood gain per added latent improves over a strong gate-count allocation baseline by 0.00553 and 0.00716 BPB on two disjoint Python test sets at matched latent spend. Repetition and local code-structure features are especially useful for ranking prefixes, even though the corresponding signals are too weak to localize individual gates. This suggests that test-time adaptation of learned tokenization is better posed first as a compute-allocation problem over prefixes, and only second as a boundary-placement problem within a prefix.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.