Label-Free Aggregation Thresholds for Set-Valued Self-Consistency
Abstract
Drawing several samples and combining them is standard test-time practice. When outputs are sets, combining means keeping the candidates that appear in more than a fraction of samples, and practice defaults to . We show that can be read off the unlabeled samples themselves. A predictor built from sampling-sharpness statistics recovers 68% of the oracle gain with no annotation, beats a grid search costing ten labelled sentences (+0.013 F1, bootstrap CI above zero), and ties one costing twenty-five; as a prior it never does worse, even refit on task families it never saw. We package it as a label-free triage rule. It works because a sample frequency is not a probability of correctness: every model family we measure is overconfident, with calibration gaps as large as -0.261 on multi-label, and the default is below optimal in 192 of 246 pairs across 5 kinds of set-valued output and 31 checkpoints. What governs the miscalibration is sampling sharpness—temperature and instruction tuning are two routes to it rather than two mechanisms, and matching on sharpness makes the apparent tuning gap vanish. One consequence inverts the usual advice: aggregation buys precision, not recall, and its gain is largest where extraction is failing, so a large gain is a cue to fix the pipeline. The threshold condition itself is classical, and a sample frequency is not the calibrated probability it assumes; we contribute the miscalibration measurement, the sharpness mechanism, and the label-free procedure. We release every sample, so all numbers recompute without a GPU.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.