acceptodds
Under review as a conference paper at ICLR 2027

MangaBench: Can Multimodal Models Reason Across Thousands of Manga Pages?

Abstract

As Multimodal Large Language Models (MLLMs) scale to context windows of up to one million tokens, many long-context benchmarks increasingly test whether models can effectively use their context. However, evaluating model understanding across long, connected visual narratives remains an open challenge, and recent benchmarks fail to evaluate model reasoning over content beyond the context window limit. To address this gap, we introduce MangaBench, a benchmark for evaluating reasoning and understanding over long visual stories. Its tasks require models to aggregate evidence and to track entities, relationships, and states across volumes. Manga is a particularly challenging medium because its information-dense panels depict selected narrative moments, and later scenes may depend on previous visual details. MangaBench pairs manually constructed questions with an evaluation of multimodal models using native agents and an RLM-inspired reader. We compare task performance across model–harness configurations. Our evaluation shows that even frontier models using these harnesses struggle with counting and listing, and with state tracking across volumes. These results highlight the need for models and agent harnesses that can more effectively gather visual evidence to reason over extended narratives.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.