acceptodds
Under review as a conference paper at ICLR 2027

Pluralistic Preference Alignment via Sortition-Weighted RLHF

Abstract

Preference-based alignment methods such as reinforcement learning from human feedback (RLHF) and Direct Preference Optimization (DPO) implicitly determine whose values a model learns through the composition of their rater pools. Because these pools are often convenience samples, the resulting preference signal may systematically overrepresent some demographic groups while marginalizing others. We introduce sortition-weighted preference learning, a representativeness-aware framework that incorporates algorithmic sortition—the sampling mechanism underlying citizens' assemblies—into preference-based fine-tuning. We study two implementations. Hard Panel trains exclusively on preferences from a quota-satisfying mini-public selected through sortition, whereas Soft Panel retains the full dataset while weighting raters according to their inclusion probabilities under the same sampling design. Using PRISM rater demographics and preference data, we fine-tune Llama models with DPO and evaluate their behavior against a 75-clause constitution elicited from a representative panel of U.S. residents. Across multiple preference-aggregation rules, Hard Panel achieves the strongest constitutional alignment, while Soft Panel consistently improves upon training on the full, unadjusted PRISM sample. Additional analyses of weighting functions, panel sizes, training gradients, and the Community Alignment dataset reveal an important data-quality tradeoff: hard selection performs best when representative subsamples retain sufficiently informative preferences, whereas soft weighting becomes more competitive when filtering would discard substantial amounts of high-quality data. Our results demonstrate that the demographic composition of preference feedback is an empirically consequential design choice and provide an explicit, auditable mechanism for aligning models with a specified target population. We also provide theoretical analysis of our framework.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.