MuST: Multilingual Span-level Machine-Generated Text Detection
Abstract
Detecting machine-generated text (MGT) is critical, as the misuse of large language models (LLMs) has raised concerns about the integrity, authenticity, and trustworthiness of information. In real-world scenarios, the task is challenging because human–AI collaboration produces hybrid texts, motivating the development of fine-grained detection of MGT. This line of research performs well on the English corpus, but its generalization to other languages is under-exploited. To address this challenge, we introduce MuST, a framework for multilingual fine-grained MGT detection. Our framework features unified modeling for two fine-grained MGT tasks, i.e., span-level detection and boundary detection. Besides, to accommodate different languages, we design augmented word embeddings by combining the information of neighboring words, N-grams, the sentence, and the normalized position. To assess our approach under multilingual settings, we further propose MuST-Mix, a fine-grained dataset for multilingual MGT detection. Evaluation shows that MuST outperforms 11 SOTA detectors in eight languages and is robust to unseen generators and unseen domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.