PuzRian: Scalable Puzzle Reasoning for Vision-Language-Action Agents
Abstract
Vision-language models (VLMs) have shown remarkable progress as interactive agents. However, systematically evaluating their long-horizon spatial reasoning remains challenging. We introduce PuzRian, a scalable Gymnasium benchmark for interactive visual-spatial reasoning, featuring 3 task families with controllable complexity, action spaces, feedback, and observability. Evaluating 9 frontier and open-source VLMs reveals strong performance on symbolic tasks but substantial degradation as spatial and perceptual complexity increases, with near-zero success on challenging 6-piece puzzles. Controlled analysis further reveals that increasing interaction budgets alone provides limited benefit, while fine-grained feedback enables more effective error correction. We further find that performance is particularly sensitive to required spatial manipulations rather than the size of the action space, and that models become increasingly brittle at recovering from imperfect intermediate states as task complexity grows. PuzRian provides a controlled testbed for diagnosing how vision-language agents plan, use feedback, and recover from errors during long-horizon interaction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.