acceptodds
Under review as a conference paper at ICLR 2027

PuzRian: Scalable Puzzle Reasoning for Vision-Language-Action Agents

Abstract

Vision-language models (VLMs) have shown remarkable progress as interactive agents. However, systematically evaluating their long-horizon spatial reasoning remains challenging. We introduce PuzRian, a scalable Gymnasium benchmark for interactive visual-spatial reasoning, featuring 3 task families with controllable complexity, action spaces, feedback, and observability. Evaluating 9 frontier and open-source VLMs reveals strong performance on symbolic tasks but substantial degradation as spatial and perceptual complexity increases, with near-zero success on challenging 6-piece puzzles. Controlled analysis further reveals that increasing interaction budgets alone provides limited benefit, while fine-grained feedback enables more effective error correction. We further find that performance is particularly sensitive to required spatial manipulations rather than the size of the action space, and that models become increasingly brittle at recovering from imperfect intermediate states as task complexity grows. PuzRian provides a controlled testbed for diagnosing how vision-language agents plan, use feedback, and recover from errors during long-horizon interaction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.