acceptodds
Under review as a conference paper at ICLR 2027

Replicable Split-Conformal Calibration in Coverage Units

Abstract

Two laboratories that calibrate the same prediction model on their own data deploy different prediction sets. The reason is that the calibrated cutoff is computed from continuous scores, so two calibrations never produce the same cutoff, even when both are correct on average. As a result, predictions stored by one laboratory cannot be reused by another, and an auditor cannot reproduce a vendor's prediction sets. We introduce replicable split-conformal calibration. It rounds each calibrated cutoff up to the nearest of a few values that both laboratories agree on in advance, so that both deploy identical prediction sets with a probability they choose, while each keeps its own coverage guarantee. The cost is larger prediction sets. How large depends on what the laboratories share to choose these values, and we prove three results about it. First, without shared information, identical prediction sets are impossible. Second, a shared random seed makes them possible, but at a cost that no seed-based threshold rule can avoid in the worst case. Third, a small public sample from the same population removes the need for a seed, lowers the cost, and gives guarantees that hold for every continuous score distribution. We implement our method on ImageNet, CIFAR-100 and four language models, GPT-2, Pythia, Qwen and Llama, and confirm the theory. On ImageNet, the chance of identical prediction sets rises from zero to above ninety percent, for prediction sets about a fifth larger. On CIFAR-100, with six times less calibration data, the same pattern holds at a higher cost. On the language models, a shared seed reaches the chosen probability of identical prediction sets in all 42 settings, and its cost falls as the calibration data grow. In a question-answering service built on Llama, two laboratories whose stored predictions differ on more than a third of the questions under standard calibration end up with identical stored predictions ninety-five percent of the time, at the cost of storing larger prediction sets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.