acceptodds
Under review as a conference paper at ICLR 2027

Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?

Abstract

Vision-language models (VLMs) are increasingly augmented with continuous or latent non-textual tokens intended to support "visual reasoning." Improved task accuracy alone does not show that models use these tokens: gains may instead arise from added context length, special-token anchoring, or training-time regularization. We formalize a diagnostic principle, Ablate-to-Validate, and instantiate it as the Token Replacement Test (TRT), a standardized suite of content-replacement ablations that measures how sensitive task accuracy is to a latent span's content while holding span placement, token budget, prompt, image and decoding fixed. We apply TRT to three off-the-shelf visual reasoning methods, Mirage, Mull-Tokens and CoVT. As a controlled testbed, we also train our own LLaVA-13B and Qwen2.5-VL-3B models to estimate relative depth by predicting and consuming continuous or discrete depth tokens. Our results show that accuracy gains can be a misleading proxy for latent-token reasoning: across all methods with continuous visual tokens, accuracy changes little when latent-token content is replaced by distribution-matched noise or shuffled across slots, and injecting an oracle yields no measurable gain, revealing a gap between "having a latent channel" and relying on its content. This gap is bolder for continuous spans: under the same test, accuracy with discrete depth tokens is more sensitive to their content on both backbones. TRT therefore shows how much and where the content of visual reasoning tokens matters for accuracy, and we recommend reporting such diagnostics as standard practice. Code will be released upon publication.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.