PVAlign: Population Value Guided Test-Time Scaling for Value Alignment
Abstract
Test-time scaling offers a promising solution for dynamically aligning Large Language Models (LLMs) with diverse human values. However, existing test-time value alignment approaches face two limitations: how to transform fragmented empirical evidence into generalizable population-level value representations, and how to prevent limited or potentially biased historical evidence from shifting model behaviors. To address these challenges, we propose \Population Value Guided Test-Time Scaling for Value Alignment (PVAlign), a test-time alignment framework that leverages population-level value representations to guide model adaptation. PVAlign first applies a product-categorical model to identify population groups from historical survey responses, constructing a Population Value Anchor that captures the population value distribution across survey. Under a constrained test-time computation budget, PVAlign further samples respondent's historical evidence. Finally, Population Value-Guided Distribution Fusion adaptively adjusts the anchor distribution by incorporating the deviation between historical evidence and population priors, with a controllable scaling factor to regulate evidence influence and mitigate potential bias. Experiments across four main benchmarks with 14 LLMs show that PVAlign improves all evaluated models on the World Values Survey and GlobalOpinionQA, and 12 of 14 models on each of WorldValuesBench and Deep Value. Additional evaluations on MultiTP and PrefEval demonstrate its generalizability to moral decisions and personalized preference following.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.