acceptodds
Under review as a conference paper at ICLR 2027

FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipeline

Abstract

As large language models (LLMs) advance in role-playing (RP) tasks, existing benchmarks quickly become obsolete due to their narrow scope, outdated interaction paradigms, and limited adaptability across diverse application scenarios. To address the benchmark limitations, we introduce FURINA-Builder, a novel multi-agent collaboration pipeline that automatically constructs fully customizable RP benchmarks at any scale. As the first benchmark builder in the RP domain designed for adaptable assessment, FURINA-Builder enables evaluation of arbitrary characters across diverse scenarios and prompt formats. From our builder, we derive FURINA-Bench, a comprehensive new role-playing benchmark generated by 20+ strong source LLMs, featuring both established and synthesized test characters, each assessed with dimension-specific evaluation criteria. Human evaluation and preliminary separability analysis justify our pipeline and benchmark design. Based on FRUINA-Bench, we discover key trade-offs between reasoning, hallucination, and RP performance across cutting-edge LLMs (general and dedicated RP models) and character types. Together, these findings demonstrate the effectiveness of FURINA-Builder and the challenge posed by FURINA-Bench, establishing a strong foundation for future research on RP evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.