acceptodds
Under review as a conference paper at ICLR 2027

Token-Level Model Blending and AI-Generated Text Detection: Separating Model Switching from the Generation Loop

Abstract

Mixing several language models within one passage is a proposed way to evade detectors of AI-generated text. Token-level blending, which draws a model at random for each short chunk, lowers zero-shot detection scores, yet it also changes the generation procedure, leaving it unclear whether detectors are vulnerable to model mixing at all. We separate the two by running every pool member alone through the identical procedure. Encouragingly for detection, switching itself adds no measurable evasion: four older open models run alone are as hard to detect as their blends (differences of at most 0.0071). When blends and single models are generated together, switching even raises detectability, by 0.0411 to 0.1049 for four newer models and 0.0271 to 0.0655 for four instruction-tuned models. At the same time, the chunked procedure itself creates a large gap: on news articles, Fast-DetectGPT AUROC is 0.7088 for the older models' blends and 0.9283 for their standard text. For the newer models, switching lowers detectability only relative to single models regenerated separately with the original procedure. We examine these patterns across three English domains and several detector families. The results suggest that model switching alone is not a route around current detectors and that evasion claims need same-pipeline single-model baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.