acceptodds
Under review as a conference paper at ICLR 2027

Question of Confidence: A Zero-Shot AI-Text Detector and Benchmark Suite

Abstract

Ask a large language model whether a person or a machine wrote a text, and it answers in words. Before it writes “human” or “machine”, the model already holds a probability distribution over the possible answers. We show that this distribution is a better detection signal than a probability the model is asked to write out, its stated probability. We propose two detectors built on it. LogJudge asks the model once and scores the text by the token probability it places on “human” versus “machine” before it answers. FusionJudge averages the token probability with the stated probability. Since a detector must rarely accuse a human, we count how many machine-written texts each detector catches while flagging at most 1% of human-written ones. Two state-of-the-art detectors, Fast-DetectGPT and Binoculars, catch 52% and 51%. With a low-cost commercial model as the judge, the stated probability catches 50%, LogJudge 65%, and FusionJudge 76%. The stated probability orders texts well, but it takes only a few rounded values, so many texts tie right where the decision is made. The token probability is fine-grained and separates them. Their average, FusionJudge, keeps the stated probability’s order, breaks its ties with the token probability, and follows the token probability where it is confident. Current detectors are near ceiling on benchmarks written by 2023-era models, so we release Manugram, about 30,000 texts that recreate a standard benchmark with four current models and add scientific abstracts, some of them written by automated AI-Scientist systems. Averaged over the four benchmarks, FusionJudge is the strongest detector we tested. Neither detector needs training, so their accuracy holds when the topic changes, while a trained detector’s accuracy collapses. They also read only the few probabilities that commercial APIs already return, which means strong API models can be used as detectors without any extra work. We release the code and data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.