SCBO: Semantically coherent batching and ordering for LLM-based social surveys
Abstract
Large Language Models (LLMs) offer a scalable way to simulate survey respondents conditioned on demographic profiles and observed reference responses. However, the conventional one-question-per-prompt paradigm is limited in three respects: it repeatedly encodes the same context, incurring substantial token and inference costs; it restricts each target to its own narrow subset of observed responses, preventing reference evidence from being shared across targets; and it predicts every answer in isolation, preventing later predictions from leveraging information in earlier answers. To address these limitations, we predict multiple questions in a single prompt, which amortizes the shared context, allows multiple targets to share a broader pool of observed reference responses, and enables later predictions to condition on earlier ones. However, such a method faces two challenges: (1) how to form semantically coherent batches and select shared references, and (2) how to order questions and references to improve autoregressive generation. We propose Semantically Coherent Batching and Ordering (SCBO), a training-free framework that addresses these challenges through two modules, preceded by a Question Instantiation step in which an LLM extracts compact semantic representations from survey items to filter out template noise. (1) Semantic Batching and Selection: SCBO groups related questions into semantic batches and constructs a shared reference bank by combining target-specific retrieval with centroid-based completion. (2) Curriculum-based Ordering: It then applies a heuristic easy-to-hard ordering to target questions and orders references within the shared bank according to their semantic alignment with the ordered questions. Experiments on four large-scale survey datasets and four LLMs show that SCBO substantially reduces token consumption and inference time while generally improving prediction accuracy over the non-batched baseline. Code is available at https://anonymous.4open.science/r/SCBO-41D8.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.