The Hidden Role of Candidate-Pool Provenance in Selector Evaluation
Abstract
Localise-then-repair systems first retrieve a small set of candidate code locations and then select the most relevant one. Existing evaluations often interpret the resulting localisation accuracy as a property of the selector, implicitly assuming that this accuracy is stable across candidate generators. We show that this assumption does not generally hold. We introduce a controlled evaluation protocol that fixes the selector, candidate-set size, gold element, and evaluation instances, while varying only the generator that supplies the distractors. Across multiple candidate generators and 7 open-weight selectors on 269 SWE-bench Verified issues, changing candidate provenance substantially shifts conditional selection accuracy, with an average gap of 8.2 points between real localisers. Despite these shifts in absolute performance, selector rankings remain largely stable. We further find that generator coverage is strongly associated with selector difficulty, enabling a simple coverage-based adjustment that improves transfer of absolute performance estimates to unseen generators. These results show that selector accuracy should be reported together with the candidate generator on which it is measured, and motivate separating retrieval coverage from conditional selection performance in localisation evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.