Simulation-Augmented Generation via Reinforcement Learning
Abstract
Simulation-augmented generation by Milli et al. (2026) is a recent framework for answering normative questions by querying generative simulations of individuals in a prompt-dependent target population and synthesizing their responses. Kraiczy et al. (2026) formalize representative selection in this setting using social choice and, to avoid simulating every individual, use proportional clustering over predicted viewpoint embeddings to route each prompt to a small subset among simulations. We introduce reinforcement learning from social choice axioms (RL-SCA), a framework for directly training representative routers using continuous measures of social choice guarantees as reward signals. In our setting, we use the mEJR+ factor, a continuous measure of proportional representation, and show that it can be computed in time, making it practical as a reward for reinforcement learning. This allows us to train a lightweight router end-to-end to map each prompt directly to a representative subset of simulations, replacing the two-stage process of predicting population-wide embeddings and subsequently clustering them. The resulting router requires only a single forward pass at inference time, avoiding predicted embedding computations for the full population while achieving significantly improved proportional representation across three datasets spanning politics, personal advice and book reviews.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.