acceptodds
Under review as a conference paper at ICLR 2027

Caught You, AI Reply! Rephrasing Messages to Expose AI-Generated Replies

Abstract

Detecting undisclosed AI-generated replies can help authors assess whether the feedback they receive reflects a respondent’s own reasoning and judgment. While existing work has primarily focused on improving detectors, we investigate a com- plementary question: Can authors make AI-generated replies to their writing eas- ier to detect by changing how they express the same content? We formulate this question as constrained optimization of source wording using the expected score assigned to subsequent AI replies by an existing AI-text detector. Rephrasing to Expose AI Replies (REAIR) provides a common framework for studying this in- tervention with BoN, Hill-climb, Hill-climb-ICL, and a source-specific adaptation of GEPA. The framework selects rephrasings using detection feedback from the replies they elicit, with semantic similarity and length checked against the origi- nal source. For greedy Gemma 4 31B replies generated after source selection, the three iterative instantiations achieve true-positive rates of 77.23% to 78.39% on an online discussion task and 56.46% to 63.27% on a research-proposal review task, compared with 29.39% and 12.93% for unmodified sources. These rates use the Binoculars AI-text detector at its fixed published threshold. Across the evaluated settings, gains are smaller under sampled generation, and transfer varies with the reply model, instruction, and detector. These findings establish source rephrasing as an author-controlled means of improving AI-reply detection that complements existing post hoc detectors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.