LiveAnyRoleBench: Towards Role-playing as Any Entity
Abstract
Large language models (LLMs) are increasingly used to simulate diverse roles, from fictional characters in games to real-world entities in social simulations. Yet existing benchmarks largely focus on few well-known characters (e.g., fictional protagonists or historical figures), who already appear in pretraining data, so even high scores may be contaminated by memorization instead of reflecting general simulation ability. Moreover, evaluating only on human characters can yield biased results, leaving performance consistency across fundamentally different entity types (e.g., warships, space missions, or animals) unclear. To address both gaps, we build role-playing tasks directly from Wikipedia, which is a continually updated source that covers an enormous range of entities. By converting each entity article into a sequence of dated actions, we introduce LiveAnyRoleBench, which evaluates whether an LLM can role-play as any entity by predicting its next action given its own history actions as context. This is, to our knowledge, the first role-playing benchmark that dates every task and spans arbitrary entity types (137 types, 25,208 entities, 969,088 actions). Evaluating 13 models, we find that every model's score drops by 3 to 18 percentage points on actions dated after its training cutoff. Across entity types, the strongest model leads almost everywhere, but among models of similar strength, rankings on human characters do not predict rankings on other types. Grounded in real-world actions, LiveAnyRoleBench provides a foundation for role-playing agents and world simulators requiring fine-grained interaction and complex decision-making. We release the dataset, construction pipeline, and all model outputs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.