acceptodds
Under review as a conference paper at ICLR 2027

Mary Had a Little [What?]: Decoding Can Matter More Than the Model Itself in Zero-Shot Detection

Abstract

Likelihood-based zero-shot detectors seek to identify the source of a text by comparing its observed tokens with what an evaluator LLM expects to see. We find that this relationship can depend less on who wrote the text than on how it was sampled, which is a routine generation choice. Across six evaluator LLMs and eight domains, applying standard truncation (top-k/top-p) to an evaluator's own generation moves its liketropy score about twice as far as switching to another LLM, and more than ten times as far as switching to a human writer. Calibrated to accept 95% of the evaluator's raw-sampled text, liketropy rejects 78% of its truncated text yet accepts 88% of human text as its own. To explain this, we derive a decomposition of the expected score for a family of log-likelihood-based detectors. It combines an entropy-based configuration term, which is highly sensitive to how narrowly the generator samples, and an identity term, which depends on the KL divergence from the generator to the evaluator. Consistent with this mechanism, in a complementary experiment involving five detectors, switching the generator from truncated to full-vocabulary sampling drags AUROCs as high as 0.99 down to chance or below, while a supervised classifier outside the family still separates the same texts with AUROC 0.84.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.