acceptodds
Under review as a conference paper at ICLR 2027

RPG-AgentBench: Evaluating the Understanding–Action Gap in Long-Horizon Role-Playing

Abstract

Large language models (LLMs) can generate convincing character dialogue, but sustained role-playing also requires decisions grounded in a character's experiences, knowledge, and obligations. However, current models can lose this grounding over extended interactions. We introduce RPG-AgentBench, a benchmark built in a role-playing game environment spanning 32 worlds and 197 characters. Agents interact through dialogue and executable actions that change the world state. Three stages first assess historical understanding and its use in decisions, then test whether these abilities remain reliable in online interaction. Here, the consequences of earlier actions and new events shape the next situation, requiring models to update their understanding of the character's circumstances and decide how to act next. Experiments across five model families show that strong historical understanding does not necessarily translate into reliable decisions. Sustained interactions further reveal that models make decisions that contradict their character's stance, established facts, or prior commitments. Even Gemini-3.8-Flash, the best-performing model in our online evaluation, maintains role consistency throughout all 40 turns in only 44.8% of evaluable pressure-test trajectories. These results reveal a capability gap between expressing a character and sustaining that role: reliable role-playing requires maintaining coherence between a character's history, current circumstances, and subsequent actions throughout ongoing interaction with a player.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.