Resolving Score-Ordering Reversal in Watermarked Text: An Asymmetric Two-Sided Test for LLM Text Detection
Abstract
LLM watermarking and LLM-generated text detection are two complementary techniques for misinformation detection, intellectual property protection, and related provenance tasks. However, we found the two conflict: the watermarked LLM text can reverse the score-ordering of the statistical score by the state-of-the-art logits-based detectors, leading to failure in LLM text detection. We dive into this phenomenon to propose a new statistical measure for LLM-generated text and design a two-sided test accordingly. The test is enhanced threefold: 1) we design the upper and lower test branches separately with a pair of asymmetric statistical measures; 2) we remove the tail of the reference distribution due to the observation that both human and LLM rarely output low-probability tokens; 3) we use the context entropy to weigh each token position differently in the statistical measure for better discrimination. On a total of 660 configuration settings (LLMs watermarking methods datasets), our detector exceeds Fast-DetectGPT and AdaDetectGPT in terms of mean AUROC by 0.1929 and 0.1986, showing higher detection rates on a wider range of LLM text.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.