acceptodds
Under review as a conference paper at ICLR 2027

LesionProbe: Detecting AI-Generated Text through Module Contribution Patterns

Abstract

The widespread adoption of large language models (LLMs) has increased the need to distinguish LLM-generated text (LGT) from human-written text (HWT). However, existing detectors often struggle under distribution shifts, limited text lengths, and paraphrasing. We investigate whether these two types of text differ in how a fixed proxy LLM's predictions depend on its internal computations, and whether such differences provide transferable detection signals. To this end, we propose LesionProbe, a framework inspired by lesion studies in neuroscience. It selectively ablates the attention and MLP branches, individually and jointly, at each layer of a frozen proxy LLM, and measures the resulting changes in next-token predictive distributions using normalized Jensen–Shannon divergence. These responses form module contribution patterns across layers and token positions. Shapley-based analysis reveals systematic layer-wise differences between HWT and LGT, while a lightweight classifier uses the extracted patterns for detection without fine-tuning the proxy LLM. Experiments on DetectRL demonstrate strong detection performance and generalization. Across generation-source configurations, LesionProbe achieves an average AUROC of and a true-positive rate of at a false-positive rate of . Under leave-one-domain-out evaluation, the corresponding results are and . Evaluations with ten proxy LLMs yield mean AUROC values of -, with further experiments showing strong performance across text lengths. Under DIPPER paraphrasing, [email protected]%FPR decreases from to . These findings support module contribution patterns as transferable detection signals and offer insight into how a proxy model responds differently to human-written and generated text.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.