PersonaAlign: Persona-Driven Interactive Alignment Benchmark
Abstract
Evaluating personalized large language models (LLMs) requires standardized interactive benchmarks, but existing benchmarks fail to capture the context-dependent, individualized nature of real-world personalization. To address this issue, we introduce **PersonaAlign**, a multi-turn interactive benchmark for learning from feedback. PersonaAlign uses _simulated users_—LLMs conditioned on personas—to create a diverse population of users with distinct backgrounds, writing styles, and feedback behaviors that define the preferences LLMs must learn. We instantiate PersonaAlign with a set of 1,000 personas with diverse writing style preferences, three writing tasks, and results from fourteen personalization methods (nine prompt-based, four fine-tuned, and a zero-shot baseline) across four different open-weight base models. PersonaAlign is modular and can be extended to new tasks, personas, and learning algorithms. Final results are reported with a pairwise ranking judge LLM. The simulated users and judges in PersonaAlign are validated through several human annotation studies. PersonaAlign strongly agrees (Spearman's ) with human rankings of four different personalization methods, differentiates between the methods, and shows there is room for algorithm improvement. We release our framework to support personalization research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.