acceptodds
Under review as a conference paper at ICLR 2027

LLM Human Response Alignment: A Multi-Sample Debiasing Framework

Abstract

Large language models (LLMs) offer scalable synthetic respondents, but their usefulness as human proxies depends on alignment with observed responses. We distinguish two sources of misalignment. Model-side bias arises when LLM and human response distributions differ under the same observed inputs. Latent-context omission limits prediction because relevant respondent information is missing. We propose a multi-sample post-hoc debiasing framework that keeps the LLMs frozen and learns an external correction from human-labeled data. The correction retains a response vector, preserving distributional variation and source identity that a single response can discard. The sampling design adapts to the task: fixed model sources for individual prediction and sampled personas for population-mean prediction. Under squared loss, the vector weakly reduces Bayes-optimal risk relative to either scalar summary, with strict gains when compression loses predictive information. Across three broad benchmarks spanning diverse domains, our method attains the best mean in every main comparison, reducing MAE relative to uncorrected LLM predictions by up to 47% at the population level and 54% at the individual level. We further show that our method preserves respondent rankings and improves revenue prediction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.