acceptodds
Under review as a conference paper at ICLR 2027

Attacking AI Editing Scores with Adversarial Rewriting

Abstract

Large language models (LLMs) now generate and revise text so rapidly that verifying its provenance has become increasingly difficult across scholarship, journalism, law, and creative writing. The written record is becoming a blend of human and machine contributions, and the central question is no longer whether AI was involved, but how much of a text it shaped and how reliably that can be measured. Models designed for this task assign AI editing scores, and because such scores increasingly inform judgments about authorship, credit, and integrity, their trustworthiness under adversarial pressure is a critical concern. In practice, however, the original human-written text is usually unavailable when a text is scored, which may make AI editing scores sensitive to surface form rather than the actual extent of AI editing. We study whether meaning-preserving rewrites can substantially reduce AI editing scores in this setting. We propose Prior-guided Rewriting for Editing Score Suppression (PRESS), which uses a paired humanization prior to identify regions characteristic of AI-edited writing and a surrogate model to select candidate rewrites. PRESS requires neither target model queries nor the original human-written text paired with the input, while a perturbation budget and quality checks control the rewriting process. The experimental results show that PRESS can lower AI editing scores across four models without moving the adversarial text closer to the original human-written text, raising concerns about their reliability as measures of the extent of AI editing.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.