Diagnosing Perceptual Bottlenecks of LLM Agents in Interactive Games with Unknown Rules
Abstract
We introduce a framework for evaluating LLM agents' perception through their action-selection reports on ARC-AGI-3, a benchmark of interactive games with unknown rules, controls, and goals. It automatically labels screen changes and uses an LLM judge to compare them with agents' reports, identifying omissions, inaccurate descriptions, and unsupported claims. Across four models, errors vary by model and change type; even explicitly reported changes are sometimes misdescribed. To test whether detected errors hinder progress, we resume execution from identical states and histories, with and without a message describing a selected unreported change. The intervention improves both reporting accuracy for the selected changes and overall level attainment in Qwen3.8-Max, highlighting the importance of accurate perception for reasoning and planning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.