acceptodds
Under review as a conference paper at ICLR 2027

VisScore: Explainable LLM-Generated Text Detection via Workspace Concept Visibility

Abstract

Detecting large language model (LLM) generated text is central to trustworthy AI, yet existing detectors, whether training-based classifiers, or statistics-based detectors, return a bare probability or score: they tell the user that a text is LLM-generated, but not which words make it look so. We start from an observation about the scoring model itself: it is an LLM, trained on text distributions similar to those of the generators it must detect, so its internal computation should resonate with the stylistic signature of LLM-generated text. Decoding this computation into the model's own vocabulary space exposes the reader's workspace and confirms the concept-visibility hypothesis: LLM-generated text saturates a small, human-readable “flowery prose” vocabulary, while human text saturates a casual, conversational register. We build on this signal: decode a single golden layer of a frozen reader, discover LLM- and human-signature vocabularies from train-split logit differentials, and score each document by how strongly its workspace activates these words. Since the score aggregates named, readable words, the evidence behind every verdict can be directly inspected and audited by a human reviewer. On DetectRL-X, spanning eight languages, four generators, and six domains, reaches document-level AUC of on English and on the 8-language mean with a Qwen3-8B reader, outperforming all 11 baselines. To further explore the potential of this signal, we fit a -parameter classifier over simple statistics of the concept-hit sequences; by recovering the structure that the scalar score discards, it lifts the mean AUC to and, at the deployment-relevant false-positive rate, the from to . also holds up in the wild: it generalizes to unseen languages, generators, and domains, needs only a few labeled examples to calibrate, and stays robust to varying text lengths, LLM-assisted editing, and paraphrasing or perturbation attacks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.