acceptodds
Under review as a conference paper at ICLR 2027

PersonaAlign: Persona-Driven Interactive Alignment Benchmark

Abstract

Evaluating personalized large language models (LLMs) requires standardized interactive benchmarks, but existing benchmarks fail to capture the context-dependent, individualized nature of real-world personalization. To address this issue, we introduce **PersonaAlign**, a multi-turn interactive benchmark for learning from feedback. PersonaAlign uses _simulated users_—LLMs conditioned on personas—to create a diverse population of users with distinct backgrounds, writing styles, and feedback behaviors that define the preferences LLMs must learn. We instantiate PersonaAlign with a set of 1,000 personas with diverse writing style preferences, three writing tasks, and results from fourteen personalization methods (nine prompt-based, four fine-tuned, and a zero-shot baseline) across four different open-weight base models. PersonaAlign is modular and can be extended to new tasks, personas, and learning algorithms. Final results are reported with a pairwise ranking judge LLM. The simulated users and judges in PersonaAlign are validated through several human annotation studies. PersonaAlign strongly agrees (Spearman's ) with human rankings of four different personalization methods, differentiates between the methods, and shows there is room for algorithm improvement. We release our framework to support personalization research.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.