The Price of Calibration in Learning from Comparisons
Abstract
We study numerical regret when each round permits at most a fixed number of comparisons and at most of rounds receive numerical audits. We first classify comparison-only learnability for arbitrary families of continuous symmetric links. Sublinear regret is possible exactly when one polynomial of degree at most the label cap recovers every numerical half-gap up to an unknown positive scale bounded away from zero. The minimax rate is then square-root; otherwise it is linear. Thus known-link reconstruction, shared numerical interpretation, and uniform signal strength are distinct requirements. We then quantify the cost of missing calibration. For a smooth one-parameter family, three labels give uniformly bounded unbiased gap estimation when calibration is known. For fixed actions and unknown parameter in a radius- interval, the minimax regret becomes . The lower bound allows audits selected after complete current transcripts. A fixed latent comparison distribution extends the same budget law to a quintic family and every fixed cap , with -dependent constants. In recorded LLM-judgment and human-score replays, a direct-judge baseline outperforms the audited methods on the principal training task; removing its rule does not reverse the reported training differences. These observations separate a worst-case need for numerical information from its finite-task value.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.