acceptodds
Under review as a conference paper at ICLR 2027

Using LLMs to Detect LLM-Generated Texts: A Cross-Generation Analysis

Abstract

Detecting LLM-generated texts is becoming increasingly important as LLMs have been widely adopted and used in various domains, while their generative outputs can also create risks of misuses and other adverse effects. Various approaches have been developed to design automated detectors, but no single solution remains reliable across models, domains, and evaluation settings. General-purpose LLMs provide an alternative and flexible way to make human/LLM authorship judgements on given texts with explanations to support their decisions, but our understanding of their detection behaviour remains limited, particularly on differences between self-detection (i.e., detection by the same text-generation LLM) and cross-detection (i.e., detection by other LLMs) and how they have evolved over time. This paper presents our study of using 15 LLMs from three generations as both text generators and detectors. We constructed 1,000 human-generated texts (HGTs) and 1,000 LLM-generated texts (LGTs) per model, and instructed each of the 15 detectors to perform binary classification tasks on all 15,000 LGTs and 1,000 HGTs, obtaining 233,428 valid judgments (excluding null/badly-formed ones) with associated natural-language explanations. Our results suggest that detection performance is more strongly related to detector generation than generator generation, while texts produced by the most recent generators remain harder to detect. Rigorous statistical comparisons of self- and cross-detection results provide mixed and model-specific results, suggesting no conclusive evidence of a consistent self-detection advantage or disadvantage across LLMs. Lastly, we analysed error patterns from three perspectives: an analysis based on false-positive and false-negative rates, an LLM-assisted analysis of LLM-generated explanations, and a literature-guided cue-based analysis. The results further show that detectors based on the first (earliest) generation of LLMs have a tendency to predict LGTs as HGTs, detectors based on the second generation of LLMs show stronger tendencies to predict HGTs as LGTs, and detectors based on the latest generation of LLMs have more balanced false-positive and false-negative rates. In addition, potential inconsistent applications of textual cues that would lead to different judgements by different LLMs are also identified and reported.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.