acceptodds
Under review as a conference paper at ICLR 2027

Amaze: Diagnosing Visual-Spatial Reasoning with Matched Maze Twins

Abstract

Vision-language models (VLMs) often fail on visual problems that require extended spatial computation, even when they can solve an equivalent symbolic problem. We introduce AMAZE , a diagnostic benchmark of procedurally generated mazes varying in scale, geometry, and visual texture. Its central design is a matched-twin protocol: each solvable maze is paired with a minimally modifiedunsolvable twin, and a model receives credit only if it classifies both correctly. This exposes a strong bias toward predicting that a path exists. To distinguish failures of visual access from failures of evidence integration, we equip agents with image-processing tools and compare monolithic execution with an orchestrated system of context-isolated specialists using local programmatic verification. Across 16 maze layouts and 1,600 matched pairs, the multi-step composition of low-level visual tools lifts paired accuracy from 12.9% to 31.6% for a monolithic agent. Orchestration with local verification further raises it to 67.6%. When a global reachability tool collapses connectivity reasoning to a single call, tools alone lift the monolithic agent to 75.8% and orchestration adds a smaller gain to 87.3%. The same pattern transfers to MentisOculi (1,000 puzzles), where accuracy rises from 58.0% without tools to 71.1% with a monolithic tool-using agent and 73.6% with orchestration. These results show that while visual tools yield large gains across visual reasoning tasks, tool access alone does not eliminate solvability bias: reliable integration of negative visual evidence is critical, and decentralization helps when that integration requires extended composition.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.