acceptodds
Under review as a conference paper at ICLR 2027

Reformulating Data-Adaptive LLM-Generated Text Detection: A Class-Conditional Gaussian Approximation

Abstract

Data-adaptive detection of LLM-generated text aims to learn from human and machine passages a scoring rule for a specified deployment criterion. We formulate the detection as a hypothesis-testing problem that maximizes machine-text detection power while controlling the false-positive rate on human writing. Existing logit-based detectors construct evidence from token log probability without considering the context, yet each token with different prefixes actually contributing differently to the detector. We therefore introduce an entropy-conditioned detection method that jointly models token log probability and contextual entropy through a two-dimensional tensor-spline witness and aggregates the resulting evidence into a passage-level score. The central challenge is learning this witness for the fixed-FPR power objective, whose direct formulation leads to an ill-posed optimization problem. We address this challenge by analyzing how token-level evidence accumulates across a passage and applying this sequential structure to approximate the human and machine scores as Gaussian distributions. This reformulation converts the original objective into a tractable problem with a regularized closed-form solution and shows that fixed-FPR power, fixed-FNR human TNR, and AUC share the same optimization direction. Across four source models and five domains, our method achieves relative AUROC-error reductions of 7.4% and 10.9% over the previous state-of-the-art method in white-box and strict proxy-model black-box evaluations, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.