acceptodds
Under review as a conference paper at ICLR 2027

CAN FACTUAL PROBES DETECT CORPUS EVIDENCE? AN AUDITABLE BENCHMARK

Abstract

Can factual probes distinguish reference-corpus evidence from answer plausibility? We introduce the Factual Exposure Benchmark (FEB), which records subject–object co-occurrences in an exhaustive scan of 6,407,814 English Wikipedia documents. Its labels describe this reference snapshot, not membership in the evaluated models’ training data. The benchmark pairs 878 evidence-absent facts with evidence-present facts matched on entity frequency and length, and includes subject-preserving object corruptions. Across five Pythia checkpoints and GPT-Neo-2.7B, pooled probe features reach mean AUC 0.5583 over five overlapping entity-grouped splits on the matched slice, compared with 0.9230 against corrupted objects. Single-model matched AUCs range from 0.5236 to 0.5676. GPT-Neo continuation-likelihood statistics also yield modest discrimination, with AUCs of 0.5488–0.5572. Pythia deduplication twins have nearly identical true-versus-corrupted answer gaps. When labels are instead defined by exact token-window presence in the released Pythia shards, the five checkpoints’ scalar answer probabilities yield AUCs of 0.501–0.518. A fixed 60-feature classifier yields 0.5545/0.5235 on all facts but 0.5026/0.4807 after exact-relation covariate matching. These contrasts show how the negative pool changes the apparent strength of factual auditing signals. They support matched evaluation of reference-corpus evidence, while leaving detection of actual training exposure unresolved.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.