Mixture-of-Experts Routing Responses Are Strong Signals for AI-Generated Text Detection
Abstract
Generalizing AI-Generated Text Detection to unseen domains, generators, and generation tasks remains challenging. We introduce a simple yet effective detector that uses the native routing responses of a frozen Mixture-of-Experts (MoE) language model. Although pretrained without any AI-Generated Text Detection objective, these routers produce responses that could be used for distinguishing authorship. Our method averages router logits over tokens, concatenates the resulting features across layers, and trains a single linear classifier while keeping the observer frozen. With GPT-OSS-20B, the classifier contains only trainable parameters and is trained on human-AI text pairs. On MIRAGE, it achieves AUROCs of , , and in cross-domain, cross-generator, and cross-task evaluations, respectively, outperforming all evaluated baselines. Controlled analyses show that detection performance is largely preserved when removing logit shifts shared across experts, but declines when expert identity is removed. Even discrete expert-selection frequencies retain a strong detection signal. Evaluations across nine frozen MoE observers further show that this routing footprint recurs across model families. These findings show that native MoE routing exposes an authorship footprint carried by relative, expert-specific preferences, and that a minimal linear readout of routing responses supports strong detection across diverse observers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.