Positional Coverage Bias: LLM Coverage Judges Verify Where Content Sits, Not Whether It Is There
Abstract
LLM judges that verify whether required content is present in a long-form report do not track content — they track position. We name and measure this failure mode of coverage-style LLM-judge metrics, Positional Coverage Bias (PCB): permuting a report's paragraphs, changing not a single word, makes the judge rescind more "present" verdicts than the agent's real content edits do — 0.108 against 0.043 patching and 0.058 regeneration on the deployed metric's own judge, Holm-confirmed, and still that judge's repeat noise ([2.1, 7.7]) when the permutation is restricted to within each section, so it is not discourse scrambling. It replicates on a second agent, two datasets and three judges under the full permutation but not under the discourse-preserving one: the effect is local, and largest where reports are weakly sectioned — the regime the audited benchmark's own agent produces, and the scope we claim. Verdicts track position far more than wording, so phantom break rates of the size deployed benchmarks attribute to agents arise with no content removed, and in our own measurements the artifact dominates the signal: a generalizability decomposition leaves the revision method only 0.6% of break-rate variance across four near-equivalent edit strategies, and a blind human pass cannot separate content-preserving controls from real loss. It afflicts every judge that reads the report and checks each item — three such protocols — and is removable by construction: in a content-determined order the judge's input is byte-identical under permutation and the break falls . The removal is exact, and held to our own threshold our construction is not the metric to adopt: it is more wording-fragile than the judge it repairs, so it relocates the nuisance instead of lowering the threshold. An existing off-the-shelf order-agnostic metric has the property already, at a threshold our own interval cannot separate from the others (+0.022 [-0.038, +0.077]). We report the negative result rather than present a dominated design as a fix. What we hand the field is the threshold: an identifiability model makes it estimable from a model-free control battery before any agent is run, below which no coverage metric can certify that content was lost, and it is non-transportable, so each benchmark must measure its own. We release the control battery, the instruments, judges and ledger.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.