acceptodds
Under review as a conference paper at ICLR 2027

CURE: Constructing User-Specific Rubrics to Evaluate Multi-Turn Role-Playing Dialogue with Agent Judges

Abstract

Role-playing dialogue is an important human–AI application across virtual companionship, narrative play, and social interaction, where users increasingly engage with characters in long-running conversations spanning many turns. These long interactions provide contexts that reveal users' evolving goals, preferences, and boundaries. They impose greater demands on role-playing models and evaluation methods: models must maintain character consistency while adapting over time, and evaluators must assess interactions against user-specific requirements. However, existing evaluation still centers largely on isolated responses or fixed single-turn probes. We therefore aim to build a multi-turn evaluation framework for role-playing dialogue. We first observe that multi-turn role-playing evaluation lacks a dedicated benchmark dataset. Using an emotional BDI mechanism, we construct and release EmoRoleBench, a benchmark of 1,751 stateful dialogues across 22 characters, 16 user profiles, and five role-playing models. Each dialogue has an independently generated post-session satisfaction label that is hidden from the evaluator. We then introduce CURE, a construct-first Agent-Judge framework that derives an explicit user-specific rubric from preferences and boundaries expressed across turns, corrects it with persona-grounded anchors, and gathers turn-level and temporal evidence before scoring. On the predefined test split, CURE achieves a Spearman correlation of 0.678 with simulated-user satisfaction, compared with 0.533 for the strongest external baseline. It also performs better in matched dialogue comparisons, remains stable across three Judge backbones, and shows positive alignment with persona-conditioned human ratings. These results show that CURE more effectively captures user preferences that emerge across turns and translates them into user-aligned evaluations of role-playing dialogue. Code and data are available at https://anonymous.4open.science/r/cure-review-artifact-2026-1E3C/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.