acceptodds
Under review as a conference paper at ICLR 2027

Stabilizing Personalized AI-Text Detection with Cross-Domain Consistent Features

Abstract

Large language models (LLMs) have rapidly advanced in generating fluent and personalized text, making it crucial to distinguish machine-generated text (MGT) from human-written text (HWT) in personalized scenarios. However, many MGT detectors exhibit significant performance fluctuations after domain transfer, as they rely heavily on features whose discriminative directions flip across domains. This motivates us to investigate whether stable discriminative signals coexist with inverted ones. Our feature-space analysis provides evidence for both an inverted component whose discriminative direction reverses across domains and a consistent component that preserves it. Building on this, we propose \method, an interpretable and detector-agnostic adjustment framework designed to counter the inversion effect. Without modifying detector architectures, \method adjusts logits using sample projections on consistent features to amplify their influence and suppress inverted ones, thereby restoring stable discrimination in personalized detection. Experiments on nine representative MGT detectors show that \method raises the mean personalized AUROC from 0.45 to 0.84 and lifts every detector above 0.79, with the same selected adjustment transferring across scenarios and across generator LLMs at a mean improvement of AUROC. \method also improves the general-domain AUROC from 0.75 to 0.88, offering a stable and interpretable solution for personalized text detection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.