LLM Content Detection and Interpretation
Abstract
In LLM-assisted writing, who develops a document's content can differ from who produces its final wording. Existing content detectors, however, can follow expression cues even when the content source remains unchanged. We study both content-source detection under changes in expression and the automatic discovery of recurring differences between human and LLM content choices across topics. To support both tasks, we decompose documents into self-contained content points and restate them in a common, plain register. Using these points as detector inputs substantially improves robustness to changes in paragraph structure and voice on HART and RAID. Across RoBERTa, RACE, and DETree, content-point inputs raise mean AUROC over 36 HART expression-change settings from 0.733 to 0.956, while avoiding the prediction reversals observed with raw-text inputs. For interpretation, we introduce content repertoire occupancy (CRO), which discovers and names recurring content choices without a predefined domain taxonomy. CRO learns a shared, overlapping repertoire of content choices without content-source labels, then compares how often human and LLM content points express each choice within matched topics. Across four writing domains, CRO identifies tens of content choices per domain, with both their population differences and descriptions supported on held-out data, revealing distinctions not made explicit in existing methods' topic words and steering descriptions. Together, these methods support a more systematic study of human–LLM content differences, combining more robust source detection across forms of expression with automatic discovery and naming of recurring content choices across topics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.