acceptodds
Under review as a conference paper at ICLR 2027

Discourse Reconstructibility in Generative Language Models: Permutation-Invariant Bipartite Infilling, Representation Boundaries, and Non-Native Linguistic Invariance

Abstract

Zero-shot detection of machine-generated text has predominantly focused on microscopic token probability curvature and cross-perplexity ratios. While computationally lightweight, token-level estimators degrade sharply under adversarial paraphrasing, lose discriminative power against frontier generative architectures, and exhibit severe algorithmic bias against non-native English (ESL) scholars (> 60% false alarm rates) due to reduced lexical entropy. In this work, we investigate generative text provenance through the lens of macro-discourse semantic reconstructibility. We formalize the proposition that autonomous generative text occupies low-entropy semantic attractor basins across multi-sentence discourse: when contiguous sentence voids are excised from discourse centroids, an independent prober conditioned on surrounding text can reconstruct omitted propositional claims with high fidelity. To resolve the permutation pathology inherent to multi-sentence discourse generation (where an infiller recovers factual claims but permutes their syntactic order), we formulate semantic alignment as a maximum-weight bipartite matching problem solved in polynomial time via the Kuhn-Munkres (Hungarian) algorithm. We conduct extensive empirical evaluations across two distinct regimes: (1) a balanced frontier multi-model cohort (N = 70) spanning GPT-4o, LLaMA-3.3-70B, Qwen-2.5-32B, and Nature STEM prose, where Hungarian Cloze achieves 82.53% AUROC [95% CI: 72.63%, 91.15%], 85.71% Precision, and only an 8.57% Human FPR; and (2) a large-scale adversarial stress test (N = 1,000) incorporating 200 DIPPER paraphrases and 100 encyclopedic rewrites. On this scaled cohort, zero-shot Cloze reveals a fundamental representation boundary (61.55% AUROC), driven by a distribution collision between adversarial paraphrasers (mean = 23.55%) and encyclopedic human science (mean = 24.65%), aligning with theoretical bounds on total variation indistinguishability. Crucially, across all evaluations, non-native ESL essays maintain low congruence (mean = 17.37%), guaranteeing an unprecedented 0.0%–2.8% false alarm rate. Our results establish the theoretical boundaries, representation mechanics, and fairness guarantees of discourse-level provenance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.