acceptodds
Under review as a conference paper at ICLR 2027

Visual Chain of Thought as a Resource-Bounded Machine

Abstract

Multimodal models now produce intermediate images and read them back through a connector. Text chain of thought theory counts only tokens, so the visual loop therefore lacks a formal account. In this paper, we model visual chain of thought as a resource bounded machine with two operators. A paint writes a pixel grid through a constant depth local decoder, and a read returns a bounded number of tokens through a connector encoder. Every bound is governed by three quantities, the pixels written at each step, the read bandwidth of the connector, and the state entropy the next step must observe. Storing one incompressible working image in text requires a quadratic number of tokens, while a single paint produces the same image in one step. Because a local decoder of constant radius implements one cellular automaton update, source to target connectivity on an open grid admits a linear number of paints, whereas a packed parent pointer transcript remains faster than linear in the diameter. When the read bandwidth falls below the state entropy, native regeneration cannot reconstruct the tape, and answers on the language side no longer depend on the image, so the intermediate image is decorative. The same bound determines what to draw, and the working image is a sufficient statistic of the state at low entropy. Frozen open weight vision language models reproduce this separation on low entropy queries. Accuracy stays at the text baseline until the connector retains enough tokens, and a corrupted render or token stream changes the answer. Larger models cross the same threshold at a smaller token budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.