MERRY: Semantic-Level Evaluation of Multimodal Emotional and Role Consistencies for Role-Playing Agents
Abstract
Multimodal Role-Playing Agents (MRPAs) are attracting increasing attention due to their ability to deliver more immersive multimodal emotional interactions. However, existing studies typically evaluate response content with textual role-playing benchmarks, while assessing multimodal expressions mainly through modality-synthesis metrics. This evaluation paradigm leaves a critical diagnostic gap: the semantics of multimodal expressions are only assessed after modality generation, making it difficult to determine whether an error originates from multimodal semantic response planning or from final modality realization. To this end, we propose MERRY, a semantic-level evaluation framework for assessing **M**ultimodal **E**motional and **R**ole consistencies of **R**ole-pla**y**ing agents. This framework extends existing evaluation dimensions to multimodal semantic responses, and introduces multimodal emotional interpretability to assess whether generated semantic descriptions support a coherent emotional interpretation. It also introduces a bidirectional evidence-finding scoring protocol for Large-Language-Model-as-a-Judge scoring, improving the alignment with human evaluators. To support evaluation, we further construct MERRY, which contains 25,507 samples with a TV-series-level separated test split. Extensive evaluations with MERRY show that current MRPAs still have substantial room for improvement in multimodal emotional and role consistency on semantic level.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.