acceptodds
Under review as a conference paper at ICLR 2027

Distribution-Level Conformal Screening for Prompt-Controlled LLM Synthetic Data

Abstract

Large language models (LLMs) can generate synthetic data at relatively low marginal cost, yet identifying prompts whose induced distributions are sufficiently aligned with a target population remains statistically challenging when only a small sample from that population is available. We cast this task as distribution-level set-valued inference and propose a conformal screening procedure that identifies, from a finite family of prompt-induced candidate distributions, those that are statistically compatible with the target population. Specifically, we construct a distribution-level nonconformity score using the empirical -divergence between real and \(N\) synthetic observations for each candidate. We then calibrate this score by its randomized conformal rank among \(m\) replicate scores generated from the candidate’s synthetic pool, and obtain the admissible subset by inverting the resulting conformal tests. Our first theoretical contribution is an exact finite-sample, candidate-wise validity guarantee. Under the compatibility null, exchangeability of the observed and calibration scores guarantees that each compatible candidate is retained with probability at least . Our second theoretical contribution is an asymptotic power guarantee: candidates separated from the target beyond a detection boundary of order , are rejected with probability tending to one. Simulations and an application to OpinionQA and MovieLens show that the retained prompts generate synthetic distributions closer to the full-data target than the empirical distribution based on the small real sample, thereby reducing downstream estimation error.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.