CATS: Power-Optimized Conformal Selection for AI-Text Detection
Abstract
Detecting machine-generated text is critical for manuscript review, academic-integrity enforcement, and content moderation, where each accusation carries direct consequences for the flagged author. However, existing methods predominantly rely on heuristic detection scores without statistical guarantees; even when calibrated to control the false positive rate (FPR), such fixed-threshold rules leave the composition of the selected set unconstrained, risking an output dominated by falsely accused human texts. In this work, we formalize machine-generated text detection as a candidate selection problem and propose Conformal AI-Text Selection (CATS), a distribution-free framework that controls the false discovery rate (FDR) in finite samples, ensuring the expected proportion of human texts within the selected set stays below a user-specified level q. Specifically, CATS estimates the selection threshold from pooled calibration and candidate data through a permutation-equivariant map and learns a score adjustment by maximizing a local surrogate for selection power. Theoretically, we prove that this pooled permutation equivariance preserves the super-uniformity of null conformal p-values, thereby guaranteeing finite-sample FDR control. Extensive experiments across ten benchmark datasets and nine detection scores demonstrate that CATS consistently maintains valid FDR control below the target level q while improving selection power over the unadjusted baseline, with an average relative gain of 38% in the main comparison.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.