CONSTRUCTING A PROVENANCE-ATTRIBUTION BENCHMARK OVER THE SCIENTIFIC LITERATURE: A SUPPLY CENSUS AND A SHORTCUT GATE
Abstract
We set out to test whether a reported operating-regime boundary in provenance-constrained generation (attribution improves when a claim could plausibly come from several sources and collapses when it could not) transfers to the scientific literature. Building the benchmark produced a different result: across three successive constructions, a rule consuming no evidence could answer it. Gold was identifiable from question surface features in 335 of 335 items; the headline metric scored 0.873 in the treatment family against 0.250 in the control at zero capability, because its denominator was the manipulation; and publication year, reachable only through the provenance map under test, identified gold in 131 of 135 items, only in the arm the map is supplied to. A suite of 639 checks passed throughout. We contribute a shortcut gate that decides which such artifacts are fatal by conditioning each candidate rule on the information it consumes: question-only rules must not beat chance, evidence-reading rules may, with their maximum published as a ceiling, and provenance-reading rules must not beat the null of their own candidate set. Applied to our construction, the gate eliminated all 27 question-level rule–scope failures and left the evidence-level ceiling at 0.994 and 0.996, which we read as a diagnosis of this design rather than an impossibility. We also report a supply census and a decision pilot of 888 calls ($10.15) that failed its prespecified validity criteria; each triggered criterion was itself misspecified, so we did not proceed to the full experimental grid.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.