acceptodds
Under review as a conference paper at ICLR 2027

WHY PERFECT AI-TEXT DETECTION IS IMPOSSIBLE: INFORMATION-THEORETIC LIMITS AND A CALI- BRATED REPORTING PROTOCOL

Abstract

We recast post-hoc AI-text detection from a moving engineering target into a quantifiable, distribution-dependent limit imposed by the entropy of natural lan- guage, and we deliver a deployment protocol that respects that limit. Our first result is a Bayes-error ceiling: any binary authorship classifier under equal pri- ors has error at least Emin = 1 2 (1 − ∥PH − PAI∥TV), so as large language models (LLMs) close the total-variation distance between their output distribu- tion PAI and the human distribution PH on the observable feature manifold, this floor rises toward 1/2 regardless of detector architecture, training data, or com- pute. Our second result is a logically independent rate–distortion floor on the paraphrase channel: measured mutual information I(H; AI) ≈ 2.71 bits per pas- sage places the channel at only ≈ 4% of its zero-information distortion floor, cer- tifying strictly positive preserved source information—a channel-fidelity quantity distinct from the marginal separability that governs detectability. We ground both bounds on a pre-registered paired corpus of 4,072 human/AI passages across six domains, generated by three open-weight models at three sampling temperatures, with a broader eight-model panel (open-weight and proprietary) for a domain- controlled robustness test. Under Benjamini–Hochberg FDR control only 8 of 12 detector/domain cells remain significant: a fine-tuned RoBERTa degrades from AUC= 0.978 globally to 0.785 on news and inverts under cross-domain trans- fer and adversarial perturbation, zero-shot logit detectors collapse to chance, and a simulated human baseline reaches only AUC= 0.542. On the two bounds we build a deployable conformal reporting protocol that issues calibrated probabili- ties with abstain regions rather than binary verdicts, with a distribution-free cov- erage guarantee. Finally, as an explicitly negative finding, we test a content/style complementarity hypothesis suggested by an uncertainty-principle analogy: the domain-pooled product K = ∆C · ∆S = 5.83×10−5 superficially matches the prediction, but an eight-model single-domain re-run reverses its sign (r = +0.41 to +0.72), exposing it as a domain-heterogeneity confound. We report this failure directly, as a methodological caution for stylometric trade-off claims.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.