acceptodds
Under review as a conference paper at ICLR 2027

Beyond Visual Token Count: Target Inclusion and Reader Recovery

Abstract

Rendering text as images can fit more source information into a model’s input budget, but token savings alone do not establish whether that information remains usable. We distinguish target inclusion—whether a queried fact reaches the reader—from conditional recovery—whether the reader can recover a fact that is present—and study both with three fixed vision-language models. On fresh held-out cases, changing image geometry improves exact identifier recovery across all three readers while preserving source content, native token counts, and input-ID sequences. In a separate finite-budget comparison on 150 fresh long-source cases, GLM answers 126 queries correctly from full-source optical input, compared with 90 from compact query-independent text. Optical input includes every target, whereas text recovers every target it retains, revealing an inclusion advantage despite imperfect optical recovery. This comparison concerns one simple static textual storage policy, not optimized textual memory. Density and dialogue-memory controls further delimit the findings. Together, these measurements show why useful compressed context depends on both inclusion and recovery, and why visual-token counts alone cannot predict task performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.