Attacking AI Editing Scores with Adversarial Rewriting
Abstract
Large language models (LLMs) now generate and revise text so rapidly that verifying its provenance has become increasingly difficult across scholarship, journalism, law, and creative writing. The written record is becoming a blend of human and machine contributions, and the central question is no longer whether AI was involved, but how much of a text it shaped and how reliably that can be measured. Models designed for this task assign AI editing scores, and because such scores increasingly inform judgments about authorship, credit, and integrity, their trustworthiness under adversarial pressure is a critical concern. In practice, however, the original human-written text is usually unavailable when a text is scored, which may make AI editing scores sensitive to surface form rather than the actual extent of AI editing. We study whether meaning-preserving rewrites can substantially reduce AI editing scores in this setting. We propose Prior-guided Rewriting for Editing Score Suppression (PRESS), which uses a paired humanization prior to identify regions characteristic of AI-edited writing and a surrogate model to select candidate rewrites. PRESS requires neither target model queries nor the original human-written text paired with the input, while a perturbation budget and quality checks control the rewriting process. The experimental results show that PRESS can lower AI editing scores across four models without moving the adversarial text closer to the original human-written text, raising concerns about their reliability as measures of the extent of AI editing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.