Tail Dependence, Not Intraclass Correlation: Design Effects and a Cluster Floor for Conformal Calibration
Abstract
Conformal prediction is often calibrated on clustered data, such as patients within hospitals or cells within donors. Clustering has little effect on marginal coverage, but it can substantially change training-conditional coverage, the coverage induced by the calibration set at hand. We give an exact finite-sample characterization of this effect. The calibration counting process satisfies a design-effect identity governed by the within-cluster correlation of threshold indicators rather than by the intraclass correlation of the scores. At extreme thresholds, this indicator correlation converges to the upper tail-dependence coefficient of the within-cluster copula. Thus two problems with the same intraclass correlation can require very different corrections: the design-effect penalty vanishes for Gaussian dependence as the miscoverage level shrinks, whereas it persists for a tail-dependent copula. The identity connects Vovk’s i.i.d. Beta law to the comonotone regime, where the effective sample size is the number of clusters. We also prove a cluster floor: no rule can guarantee PAC coverage, and no procedure can guarantee conformal risk control, from fewer than clusters; for , this is 22 clusters. Finally, we construct dependence for which a single-number summary of within-cluster correlation understates the design effect by a factor approaching the cluster size. On multi-site neuroimaging, threshold-level calibration is the only implementable method among those compared that meets the empirical PAC target.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.