acceptodds
Under review as a conference paper at ICLR 2027

From Pixels to Symbols: Machine-Native Symbolic World Models for Multimodal Reasoning

Abstract

Multimodal large language models (MLLMs) are often trained and evaluated with visual inputs designed for human perception rather than machine reasoning. Inspired by the reverse of graphical user interface (GUI), we propose ASCII maps as a compact, discrete, symbolic world model that convert visual scenes into layouts encoding walls, free space, objects, targets, and action-relevant entities. To evaluate the performance of ASCII map as a symbolic visual representation, we introduce SymbolicBench, which covers short-range path planning, long-horizon navigation, and object-centered action planning across maze, floorplan, and agent environments using grid-like images. Across top-tier proprietary and open-source MLLMs, ASCII-based inputs substantially outperform image-only inputs, achieving from 2 to 83 times improvement, while reducing token usage. Further analysis shows that these gains persist across different task complexity levels and generalize beyond grid-like images to real-world images for spatial reasoning tasks with strong robustness. variations. Together, these findings support compact symbolic world representations as an effective interface between visual scene structure and MLLM reasoning, enabling more reliable spatial reasoning and planning with fewer input tokens. The source code will be fully released upon paper acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.