acceptodds
Under review as a conference paper at ICLR 2027

Generalizing Machine-Generated Text Detection to Unseen Generators via Harmful Subspace Removal

Abstract

Machine-generated text (MGT) detection is essential for identifying the provenance of online content, yet existing detectors often generalize poorly to unseen generators. We find that this failure is closely tied to generator-specific shortcuts: detectors trained for human–AI classification encode strong generator identity information, and their cross-generator transfer patterns are highly structured rather than uniformly transferable. A natural solution is to identify and remove generator-cue tokens, but our analysis shows that token-level deletion is ineffective because cue-bearing tokens also contain useful human–AI discrimination evidence. Motivated by this observation, we propose Harmful Cue sUbspace Removal with REplay Pool (CURE), a representation-level purification framework for robust MGT detection under unseen-generator shifts. CURE first uses a frozen generator-ID teacher to estimate token-level generator cues, then maintains a three-stream replay pool to construct a low-rank harmful cue subspace by contrasting harmful-cue and clean-reference representations. Instead of dropping tokens, CURE selectively removes generator-specific directions inside token hidden states with confidence-guided strength, while training the detector on both original and purified views. Experiments on 12 unseen generators show that CURE achieves an F1 score of 89.9, outperforming 8 representative detectors by 8.5 points. Further analyses demonstrate that subspace-level removal is substantially more effective than token-level deletion, suppresses generator-specific information while preserving human–AI discrimination signals. We will release our code and datasets upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.