Evaluating Text Segmentation for Representation Quality in a Low-Resource Language
Abstract
Text segmentation can directly affect the quality of downstream language representations, particularly in low-resource, non-delimited languages such as Myanmar, where syllable and word boundaries do not necessarily align. We evaluate five segmentation approaches—grammar-informed rules, dictionary lattices, statistical CRFs, and two neural language models—on 1,100 verified sentences across Government, News, and Novel domains. Beyond intrinsic segmentation performance, we assess downstream effects using three multilingual subword tokenizers and three embedding models. Grammar-informed segmentation remains stable across domains (Boundary F1: 0.771–0.815), while neural language models show greater degradation on literary text, with Boundary F1 decreasing by up to 0.168. Boundary accuracy is strongly correlated with semantic embedding quality (Pearson , ), while poor segmentation increases subword sequence length by up to 20%. These results show that segmentation quality is an important factor in downstream representation quality and that explicit linguistic knowledge remains valuable for robust low-resource language processing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.