MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents
Abstract
Agentic perception requires agents in partially observable environments to coordinate active sensing with working memory to construct and maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this capability because they couple active vision with complex embodied navigation and control dynamics. To isolate this capability, we introduce MNIST-PRO, which converts digit recognition into a sequential search task with controlled access to past glimpses. Evaluating ten multimodal models across in-context memory representations reveals that accuracy drops from up to 99.0% on unmasked images to between 7.0% and 43.5% under partial observability. We identify three bottlenecks in the perception loop. First, agents often fail to search thoroughly, making premature guesses based on partial strokes instead of fully utilizing their exploration budget. Second, even when agents uncover most digit strokes, they struggle to combine separate glimpses in memory, a limitation that visual canvas reconstruction overcomes by recovering over 50 percentage points on two-digit tasks. Third, even with near-complete coverage on an assembled canvas, agents still make visual interpretation errors, showing that downstream recognition remains an independent hurdle. Additionally, equipping agents with persistent memory in a harness fails to improve accuracy, instead causing them to waste steps repeatedly re-checking strokes they have already seen.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.