From aggregate scores to cross-probe claims: Auditing knowledge-editing rankings
Abstract
Knowledge-editing benchmarks rank editors by averaging scores over several kinds of test question, or probes. A higher average shows one editor is ahead on the benchmark's chosen mix, not on every probe: when editors trade wins across probes, some reweighting reverses their ranking. Over two released benchmarks and a preregistered factorial we ask how often that happens and what the disagreements track: cross-probe disagreements are common, most fail at least one check tied to how the evaluation was built or scored, and the rest divide into real editor failures whose size the design cannot calibrate and a small residue no check here touches. One pair shows what such an average is made of: HalluEditBench's Generalization Score puts ROME 20.3 points above LoRA on Mistral-7B, yet LoRA leads by 21.2 on the rephrase question, because after LoRA's edit the model usually no longer answers yes/no questions in the required form (13.2% format compliance against ROME's 93.4%, from the same unedited model). Across that benchmark 12 of 17 rejections disappear without its open-answer rephrase questions†; on KnowEdit, one cell scores probe types on disjoint items, and two paired reversals survive every check the release allows. In the factorial, three large reversals between likelihood-ranked choice and token-F1 generation disappear under four other scores, including the teacher-forced accuracy KnowEdit applies. But which score shows a reversal is not stable across model and configuration: at EasyEdit's own FT-L configuration one survives two of three, and on the release's own backbone teacher-forced accuracy yields a corrected reversal where token F1 yields none. Splitting KnowEdit's pooled Reasoning label yields reversals real in sign whose magnitude a ceiling-against-chance comparison cannot calibrate. We release an audit reporting, per pair, the verdict, how far the ranking is from flipping, the largest reversal the data cannot exclude, and which confounds apply.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.