acceptodds
Under review as a conference paper at ICLR 2027

How Synthetic Labels Improve Conformal Prediction: A Perspective on Conditional Coverage

Abstract

Conformal prediction provides distribution-free, finite-sample marginal coverage, but post-hoc calibration data may be too limited to learn how uncertainty varies across inputs. At the same time, abundant covariates can often be labeled cheaply by existing domain models or general-purpose large language models. We study whether these synthetic labels can improve conditional coverage when only a small trusted sample is available. Building on the connection between conditional coverage and score-quantile regression, we introduce prediction-powered quantile learning: a large synthetic-labeled sample estimates pinball risk, a paired trusted sample corrects its bias, and an independent trusted split performs final conformalization. The final conformal correction changes which learning gains survive deployment. It exactly removes global threshold shifts, locally projects the remaining shape error through the boundary density, and leads to a three-resource expansion and a benefit–cost rule for synthetic power. Across eight classical regression benchmarks, synthetic-powered quantile learning substantially improves downstream conditional coverage while preserving marginal validity and producing more compact prediction sets. A realistic human-rating study further shows that labels from general-purpose language models improve final conditional coverage under a practical quality–quantity–cost tradeoff.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.