Detection of Machine-Generated Text Using Fine-Tuned Transformers and a Compact Classifier Based on Token-Level Probabilistic Features
Abstract
The widespread adoption of large language models (LLMs) has intensified the need for reliable detection of machine-generated content from human-written text. This work proposes and evaluates two complementary detection strategies: (1) fine-tuned pretrained transformers (RoBERTa, Llama 3.2 3B, and Qwen2.5-1.5B) for text classification, and (2) a compact sequence classifier based on token-level probabilistic features. Both approaches are trained on an extensive corpus combining the M4 and CUDRT datasets, which contain human-written and machine-generated texts. We evaluate the models on a diverse set of challenging benchmarks, including outputs from multiple LLMs and texts edited by humans or LLMs. Across five held-out benchmarks, detector performance was strongly dataset-dependent, and no single approach dominated across all evaluation settings. The Feature Model achieved the strongest performance among our individual models on selected benchmarks, while fine-tuned transformers performed better on others. A hard-voting ensemble combining the Feature Model with two fine-tuned transformers achieved the strongest result on BUST among the evaluated detectors. Ablation experiments show how feature types and feature-generating models affect detection performance. Using the same probabilistic inputs, the Feature Model achieves higher Macro-F1 than Logistic Regression and XGBoost baselines on all five held-out benchmarks. The Feature Model itself contains approximately 100k trainable parameters, although the computational cost of end-to-end inference is dominated by the four LLM feature generators. Overall, the results demonstrate the complementary strengths of fine-tuned and probabilistic-feature detectors rather than the universal superiority of either approach.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.