acceptodds
Under review as a conference paper at ICLR 2027

Expressivity, Topology and Routing in Latent Visual Reasoning

Abstract

Autoregressive Large Language Models (LLMs) inherently lack grounding in the physical world and are fundamentally bounded by text statistics. While multimodal distillation allows models to internalize a visual "mind's eye" via continuous latent tokens without requiring inference-time images, it introduces critical confounders regarding computational expressivity, geometric topology, and causal routing. In this work, we introduce Latent Visual Reasoning (LVR) and show that distilling the continuous CLIP manifold into an autoregressive prefix improves performance on machine translation tasks. To mechanistically isolate and probe the underlying representational dynamics, we formalize LVR as a controlled distillation framework equipped with a Warmup-Stable-Decay (WSD) projection head. First, to isolate **computational expressivity**, we disentangle semantic grounding from structural algorithmic buffering, proving that continuous visual hypersphere alignment injects strictly positive mutual information rather than merely exploiting extra compute steps. Second, examining **geometric topology**, we demonstrate LVR induces a mathematically robust -Interlingua. Despite English-only projection head fine-tuning, it maps zero-shot French and Czech inputs to highly localized geometric coordinates. This establishes that the mapping from multiple human languages to the visual modality serves as a universal geometric interlingua, demonstrating that the representation captures physical concepts rather than merely memorizing English syntax. Third, to evaluate **causal routing**, we conduct a controlled ablation study to compare the causal routing capabilities of modular cross-attention architectures with those of early-fusion models. We reveal that while modular LLMs encode visual geometry, their decoders suffer from causal attention collapse, providing strong mechanistic evidence that effective routing relies on the self-attention circuits of a natively unified foundation. Our results highlight the limitations of grafting continuous visual bottlenecks onto modular architectures, pointing toward a broader directive: future reasoning frameworks should prioritize natively unified foundations to treat any continuous non-linguistic manifold as a fundamental computational primitive.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.