acceptodds
Under review as a conference paper at ICLR 2027

Local Entropy Fluctuations for Black-Box LLM-Generated Text Detection

Abstract

Zero-shot detection of machine-generated text is vital for mitigating the risks of large language models (LLMs) misuse. While state-of-the-art dual-model and resampling-based detectors outperform single-model statistical detectors on unprompted continuations, they suffer a catastrophic breakdown on instruction-conditioned text, especially in the black-box setting where the source model is unknown. We analyze this failure and trace it to poor prompt conditioning for resampling-based methods and diminished cross-model discriminability for dual-model detectors. Motivated by recent study that human writing maintains a steady, uniform information flow, whereas machine-generated text exhibits sharp, localized entropy fluctuations, we propose a new metric, the Relative Peak Prominence Rate (RelPPR) to capture this structural contrast along the entropy trajectory and further leverage it as a signal to for LLM-generated text detection. Extensive evaluations across diverse model families and datasets demonstrate that RelPPR achieves state-of-the-art detection performance in the black-box setting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.