BiasFate: Calibration Limits of Reweighted LLM Simulations
Abstract
Can reweighting LLM-generated survey responses improve on the human sample used to calibrate them? BiasFate links calibration information, proposal coverage, and estimation error for categorical responses. With reused empirical proposal counts and no clipping, self-normalized importance sampling (SNIS) equals the calibration mean restricted to observed categories; full empirical coverage recovers the direct mean exactly. Under covered support with known proposal probabilities and independent draws, unnormalized plug-in weighting adds nonnegative variance, while SNIS inherits calibration variance to leading order. Support truncation and clipping introduce distinct sources of bias. Across roughly 640K generations, four –B models, a 34B replication, four task families, and released GlobalOpinionQA marginals, the direct calibration mean beats deployed clipped SNIS in 270 of 288 matched configurations. Removing clipping reduces covered-cell SNIS RMSE from 0.1027 to 0.0245, versus 0.0148 for direct estimation. Plateau diagnostics transfer to held-out questions, while diagnostic-guided routing and allocation show no improvement over matched controls. These results explain when simulation reproduces calibration information, how it loses target mass, and why diagnostic accuracy and estimation gains require separate evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.