AnchorWrite: Detector-Query-Free Rewriting of LLM-Generated Text via Human-Reference Alignment
Abstract
Large language models (LLMs) generate fluent text that can remain statistically distinguishable from human writing. Existing post-generation rewriting approaches are often evaluated primarily through detector reduction, making it difficult to separate detector-specific optimization from movement toward an independently defined human linguistic reference. We introduce AnchorWrite, a detector-query-free framework that treats rewriting as constrained alignment to a fixed human reference. AnchorWrite constructs cross-order language-model statistics from human-authored text, calibrates them against a length-conditioned human reference, and uses token-level evidence to localize regions of linguistic deviation. A pretrained language model proposes local rewrites, which are accepted only when they reduce human-reference distance while satisfying semantic, linguistic, and edit constraints. Across six domains, mean human-reference distance reduction increases from 0.068 at 5.60% realized editing to 0.363 at 24.32%. On frozen outputs, AUROC decreases by 0.057-0.089 across six external AI-text detection methods, despite no detector feedback during rewriting. Matched-edit comparisons with instruction-based humanization and neural paraphrasing show that detector responses remain transformation-dependent. At the strongest regime, BERTScore remains 0.968 while perplexity increases from 8.88 to 10.71. These results position human-reference alignment as an interpretable control signal for constrained rewriting, with reduced detector separability emerging as an external transfer effect rather than an optimization target.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.