From Pixels to Symbols: Machine-Native Symbolic World Models for Multimodal Reasoning
Abstract
Multimodal large language models (MLLMs) are often trained and evaluated with visual inputs designed for human perception rather than machine reasoning. Inspired by the reverse of graphical user interface (GUI), we propose ASCII maps as a compact, discrete, symbolic world model that convert visual scenes into layouts encoding walls, free space, objects, targets, and action-relevant entities. To evaluate the performance of ASCII map as a symbolic visual representation, we introduce SymbolicBench, which covers short-range path planning, long-horizon navigation, and object-centered action planning across maze, floorplan, and agent environments using grid-like images. Across top-tier proprietary and open-source MLLMs, ASCII-based inputs substantially outperform image-only inputs, achieving from 2 to 83 times improvement, while reducing token usage. Further analysis shows that these gains persist across different task complexity levels and generalize beyond grid-like images to real-world images for spatial reasoning tasks with strong robustness. variations. Together, these findings support compact symbolic world representations as an effective interface between visual scene structure and MLLM reasoning, enabling more reliable spatial reasoning and planning with fewer input tokens. The source code will be fully released upon paper acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.