acceptodds
Under review as a conference paper at ICLR 2027

Trust, but Verify: A frontier model benchmark and a leakage-free protocol for AI-generated text detection

Abstract

AI-generated-text detectors increasingly inform academic-integrity and content-moderation decisions, yet reported accuracy is hard to trust. We audit three vulnerabilities. (1) Evaluation leakage: test-informed checkpoint selection inflates macro-F1 by 27.3 points in a controlled MAGE replication. Our COLING-2025 endpoints differ by 1.72 points (83.65 → 85.37), but incomplete candidate histories prevent attributing that difference solely to selection. (2) Distribution-sensitive human-text flagging: on domain-matched scientific abstracts, RoBERTa-large-OpenAI flags 45.5% of post-2023 text against 20.1% of pre-2015 text. This direction persists within both PubMed and arXiv, although source composition changes some detectors’ aggregate conclusions. (3) Heterogeneous frontier performance: evaluated off-the-shelf detectors span macro-F1 58–85 on 2026 frontier text. The author-released DetectAnyLLM checkpoint achieves the highest external macro-F1, 84.60, compared with 76.67 for the COLING-2025 winner Advacheck, yet flags 48.78% of our post-2023 presumed human abstract probe. We evaluate the author-released checkpoint, not a reproduction of the published result. Rankings change with the operating point, and Binoculars and Fast-DetectGPT recover only 14–19% of Claude Sonnet 4.6 and Grok-4.20-reasoning text. We release FrontierBench-2026, comprising 216,352 texts from 16 collected frontier LLM variants (15 evaluated), a provenance-verified pre-2015 human reference, a domain-matched contemporary probe, and an evaluation harness separating development-set selection from test reporting. Using a graph-fusion detector as an audit instrument, we find no stable advantage over the shared-task winner across three seeds (paired macro-F1 difference +0.08 ± 0.24, mean ± s.d.). Frontier-aware training raises frontier macro-F1 from 78.4 to 92.8 on the shipped split and achieves 89.8 ± 0.2 on a content-controlled split with disjoint prompts and parent documents. Held-out-provider recall is 0.845 shipped and 0.764 content-controlled. Capability-tier transfer is provider-asymmetric, while short-form detection and genre-dependent contemporary-human false alarms remain unresolved.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.