PluralisticDataSmith: Turning Pluralistic Alignment Goals into Synthetic Data at Scale
Abstract
Current alignment methods optimize language models toward a single set of values, failing to capture diverse preferences for how models should reason, communicate, and act. Yet the field lacks both the training data and benchmarks needed to reliably train and evaluate steerable value alignment. We present PluralisticDataSmith, a scalable synthetic data generation framework that produces pluralistically aligned training data. A multi-stage LLM-ensemble agent generates approximately 200 fine-grained constitutions (e.g., ensuring correct honorific usage) for each of 243 value specifications, ranging from Socratic questioning and Korean society to composite applications such as debate preparation combining formal logic and persuasive storytelling. Using a frontier-class open-source LLM (2.8T parameters) as the synthetic data generator, our pipeline produces 1.2M synthetic QA pairs across nine categories (communication, creativity, culture, philosophy, politics, professional practice, religion, wellbeing, and composite applications) with quality competitive with closed-source alternatives. An iterative refinement loop improves data quality through intuitive natural-language feedback. We further introduce PluralValueEval, a retrieval-augmented, evolving-rubrics benchmark validated by 324 human annotators across 36 specifications. Evaluation rubrics are evolved through strong–weak model contrast to assess genuine value alignment rather than compliance with artificial restrictions, with meta-judge filtering retaining 1,976 high-quality criteria. Even frontier LLMs score below 50% on our benchmark. Fine-tuning open-source models (9B–31B) on the corpus with our Grouped Robustness Optimization (GRO) yields downstream gains of 8% over base models while reducing the share of degraded specifications from 50% to 6% and the largest per-specification drop from -16% to only -1.4%. Built entirely on open-source models, our framework offers an accessible alternative to closed-source APIs. We release the training corpus, pipeline, and evaluation benchmark to support reproducible research in pluralistic alignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.