acceptodds
Under review as a conference paper at ICLR 2027

When and Where Can Learned Tokenization Be Refined at Test Time?

Abstract

Learned byte-level tokenizers expose an inference-time degree of freedom that fixed subword tokenizers do not: the model can vary how much of its learned latent resolution is used for an observed prefix. We study two forms of test-time refinement—where to restore boundaries inside a reduced native mesh and when to return a prefix to full native resolution. The experiments reveal a clear granularity asymmetry. Fine-grained boundary placement is unstable: dense insertion hurts and first-layer geometry or attention does not yield a reliable long-horizon location rule. Prefix-level allocation is substantially more predictable. A lightweight router that estimates continuation-likelihood gain per added latent improves over a strong gate-count allocation baseline by 0.00553 and 0.00716 BPB on two disjoint Python test sets at matched latent spend. Repetition and local code-structure features are especially useful for ranking prefixes, even though the corresponding signals are too weak to localize individual gates. This suggests that test-time adaptation of learned tokenization is better posed first as a compute-allocation problem over prefixes, and only second as a boundary-placement problem within a prefix.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.