COVERAGE IS NOT CONVERSION: A SAME-POOL AUDIT OF LLM SELECTION
Abstract
An LLM may generate a correct answer yet fail to select it. Consider a gate that switches between a baseline and a proposed selector. Even a perfect gate cannot fix cases where both actions miss a correct candidate. We develop a same-pool audit that compares this missed opportunity with the most a gate could add to the always-deployed proposal. It also separates information loss from errors in using that information. One joint confidence set answers two questions: does the proposal improve the baseline, and does missed opportunity exceed the remaining gate budget? The criteria are sharp for that set. In a secondary analysis of nine fixed LiveAoPS and FEVEROUS comparisons, the budget ordering holds even though no positive gain is certified. A prospective MedMCQA study improves answer accuracy by 1.005 percentage points over direct answering on 20,000 new groups, but loses .470 points to unchanged plurality. A further 4,494-group reasoning study does not confirm improvement over ordinary reasoning. The audit thus identifies limits on gate-only improvement; a proposed intervention still needs validation against a strong control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.