ASD: Auditing Source Dependence in LLM Judges For Scientific Benchmarking
Abstract
LLM judges are widely used to evaluate AI agents on open-ended scientific benchmarks, where expert grading is prohibitively costly. These benchmarks are often constructed from published papers, which raises the question of whether a verdict reflects the quality of the answer or access to the source paper, since judges can often consult the same literature. We propose Auditing Source Dependence (ASD), a two-stage pipeline that measures source dependence as the change in a judge's verdict when the source content in its evidence is varied. ASD first builds a catalog of relevant and irrelevant papers for each task and splits the source paper into passages (paragraphs and topic blocks) labeled by rhetorical role. It then judges the same answer while swapping the source paper in and out of the judge's evidence. To locate source dependence within the paper, it further applies counterfactual edits to single passages of the source paper. Across BiomniBench, BixBench, and LABBench2, replacing a relevant paper with the source paper raises judge accuracy, and editing a single passage can reverse the verdict on an unchanged answer. Methods and results passages account for over 60% of the most influential passages, and edits to different passages shift verdicts in opposite directions, so small average effects can mask large passage-level effects. These findings motivate reporting the source sensitivity of LLM judges alongside agent performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.