acceptodds
Under review as a conference paper at ICLR 2027

Simulating Experiments for Method Testing

Abstract

Field experiments (A/B tests) provide credible benchmarks for methods in societal systems, but their cost and latency slow method development. Synthetic evaluators, e.g., LLM-persona panels, offer a scalable alternative, raising the question of when they preserve the benchmark interface that adaptive methods optimize against. We show that replacing human evaluation with synthetic evaluation is indistinguishable at this interface from changing only the evaluation population if and only if (i) methods observe only aggregate outcomes (aggregate-only observation) and (ii) feedback depends on the submitted artifact rather than the method's identity or provenance (method-blind evaluation). LLM-persona panels instantiate this framework by sampling profiles and using conditioned LLMs to generate evaluations. We then separate validity from usefulness: an information-theoretic discriminability measure yields explicit bounds on the effective number of independent evaluation units needed to distinguish meaningfully different methods at a chosen resolution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.