Reverse ToBlend: Analyzing Detection Signals under Token-Level Multi-Model Blending
Abstract
Detectors and attribution tools for AI-generated text usually assume one model per passage; token-level blending breaks this by drawing each short continuation from a different model in a pool. What a defender can recover from such text remains open, even with every pool model available: machine generation, blending, or the model behind each word. We study this case with likelihood statistics from every pool model, supervised text classifiers, and 600 new continuations whose generating model is logged at every step. Encouragingly, detection is reliable: a classifier over pool likelihoods reaches AUROC of at least 0.85 in all 36 pool, domain, and generation settings, and a fine-tuned RoBERTa classifier is stronger on one of two pools. Both complement an off-the-shelf detector, RADAR, in all 18 settings where it applies. At the same time, how text was blended is hard to recover: on the logged continuations, blended and single-model text separate with AUROC of at most 0.58, and word attribution reaches 41.3% (25% chance) even with the logged decoding distributions. We examine two four-model pools across three English domains. Detection and source tracing thus appear to be distinct problems: detection alone says little about which models wrote a passage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.