SyntaxBPE: Learning Which Code Boundaries to Cross
Abstract
Most code tokenisers derive BPE merge rules primarily from corpus co-occurrence statistics, with limited use of the explicit structure of source code. This leads to a trade-off: frequency-driven merging may cross meaningful program boundaries, while rigid structural constraints can block useful recurring code patterns from being merged. This raises a question: which structural boundaries should constrain tokenisation, and which should remain open to merging? We study the coupled design of boundary policies and BPE vocabularies, and introduce SyntaxBPE, a two-stage tokeniser construction procedure. First, SyntaxBPE groups recurring boundary patterns into deployable operators. A shared model trained on randomised boundary views provides paired code-length estimates, which are used to select a deterministic policy through a lower-confidence-bound rule. Second, program structure guides vocabulary induction by downweighting merge occurrences that cross the remaining parser-visible boundaries. Structural information is used only during construction, leaving the final tokeniser deterministic and parser-free at deployment. Experiments with 1B and 3B code language models show consistent improvements over GPT-style tokenisers under both fixed-token-budget and matched-raw-data comparisons, with gains persisting after instruction tuning. These findings support using program structure as soft guidance while allowing selected cross-boundary patterns.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.