WINMORPH: DOES TOKENIZING CODE IDENTIFIERS AT MORPHEME BOUNDARIES IMPROVE LANGUAGE MODEL PERFORMANCE ?
Abstract
Code identifiers such as AddAccessAllowedAce combine meaningful sub-parts (morphemes), such as Add, Access, Allowed, and Ace. However, a language model observes only tokenizer outputs, and a tokenizer may place boundaries that do not match these morphemes. Although better boundary placement is often assumed to help, it is difficult to isolate because standard tokenizers confound alignment with token count; making more correct splits typically also changes the number of tokens. We therefore ask: when the token count is held fixed, does aligning token boundaries with reference morpheme boundaries improve language modeling? We address this question with W IN M ORPH , a corpus of 83,900 Windows identifiers with gold cut points derived from published naming rules, and a benchmark that varies the fraction of correct cuts from none to all while keeping tokens-per-identifier constant. We study two controls: (i) segmentation matching under a shared vocabulary cap, and (ii) matching the average number of model-input symbols on the training set by adjusting vocabulary coverage. On held-out identifiers, full alignment reduces bits per character by 39.8% under segmentation matching and by 18.9% under training-symbol matching, compared to no reference-aligned cuts. These results show that, when we control for token count, splitting identifiers at reference morpheme boundaries improves language model performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.