EXPOSURE EMBEDDINGS FOR CORPUS AUDITING AND DOWNSTREAM EVALUATION
Abstract
Author-conditioned title likelihood offers a behavioral signal for corpus auditing, but its interpretation depends on document identity and the construction of comparison papers. We study this signal using Comma v0.1 and the arXiv component of Common Pile. Starting from metadata-matched paper pairs, an independent identifier audit retains 1,222 pairs whose positive paper is present and negative paper is absent from that component. On these pairs, Comma’s title score achieves AUC 0.6523 with matched distractors and 0.6139 with lexically similar distractors, compared with 0.5316 and 0.5154 for GPT-Neo-2.7B. These associations concern component presence, not verified membership in the complete training mixture. Replacing the correct author list with a within-year, one-to-one derangement reduces Comma’s AUC by 0.0369 and 0.0337 under the two distractor conditions; the paired intervals exclude zero, while GPT-Neo changes by less than 0.0006. Comma’s exploratory within-year AUCs range from 0.642 to 0.663 in 2020–2023, compared with 0.455–0.569 for GPT-Neo. In a separate citation-prediction cohort, holding out query papers gives SPECTER2 AUC 0.90756 and fusion AUC 0.90770; the paired 90% interval for the difference is [−0.00095, 0.00119]. The results distinguish behavioral discrimination of an audited corpus component from incremental utility in one downstream task.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.