TOKENPROBE: What Token Counts Reveal About Black-Box Tokenizers
Abstract
Token-count interfaces reveal how many tokens a tokenizer emits, but not necessarily where the token boundaries lie. We construct two byte-level BPE tokenizers that share a reachable vocabulary and agree on the count of every input, yet segment some inputs differently; no number of adaptive count queries can distinguish them. Independent copies yield exponentially many count-equivalent segmentation functions, and a distributional lower bound quantifies the error of any point reconstruction. We then characterize what a transcript does identify: a task output is determined exactly when all consistent tokenizers agree on it. TOKENPROBE operationalizes this characterization through finite-class elimination, driven by standard active-query objectives, and task-specific certification. Exhaustive enumeration verifies the construction on all nonempty strings over of length at most ten. In a fully enumerated family of 24 BPE tokenizers with uniformly weighted targets, identifying the count-response class yields perfect held-out counts, yet a fixed representative correctly segments only of target–input pairs. Segmentation unanimity certifies of pairs without error, and boundary-level certificates still determine of internal byte positions on the remaining pairs. These results establish a sharp separation between estimating token counts and recovering a tokenizer's segmentation behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.