acceptodds
Under review as a conference paper at ICLR 2027

Evaluating Text Segmentation for Representation Quality in a Low-Resource Language

Abstract

Text segmentation can directly affect the quality of downstream language representations, particularly in low-resource, non-delimited languages such as Myanmar, where syllable and word boundaries do not necessarily align. We evaluate five segmentation approaches—grammar-informed rules, dictionary lattices, statistical CRFs, and two neural language models—on 1,100 verified sentences across Government, News, and Novel domains. Beyond intrinsic segmentation performance, we assess downstream effects using three multilingual subword tokenizers and three embedding models. Grammar-informed segmentation remains stable across domains (Boundary F1: 0.771–0.815), while neural language models show greater degradation on literary text, with Boundary F1 decreasing by up to 0.168. Boundary accuracy is strongly correlated with semantic embedding quality (Pearson , ), while poor segmentation increases subword sequence length by up to 20%. These results show that segmentation quality is an important factor in downstream representation quality and that explicit linguistic knowledge remains valuable for robust low-resource language processing.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.