Baseline Adaptation Changes Attribution in Fixed-Answer Selective QA
Abstract
We study how adapting a confidence baseline changes the attribution of gains in selective question answering. Each question's original answer is fixed, so scorers change only which answers are retained. At half coverage, native confidence refitting improves reference EM by about 3.5 percentage points in two cohorts with no overlapping normalized questions. The first is a frozen seven-scorer comparison; the second is a separately reported native-only evaluation with policies frozen before generation, with comparison-feature evaluation unfinished. Fusion loses its apparent advantage over historical confidence when compared with the adapted practical baseline, while its increment over a joint-target-matched control remains uncertain. Matched-count controls and intercept interventions separate sample count, score offsets and ranking changes. Direct comparison features show exploratory signals that depend on coverage. Native confidence already contains contextual information; the question is the value of additional comparison information. Decision accounting and case review connect measured gains to retained answers while distinguishing reference matching from factual correctness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.