acceptodds
Under review as a conference paper at ICLR 2027

Unifying Heterogeneous Visual Reasoning in Expressive Latent Space

Abstract

Visual reasoning increasingly relies on heterogeneous intermediate representations, including textual rationales, regions, scene graphs, segmentation, depth, and visual manipulations. Because these representations differ substantially in structure and modality, they are typically modeled with separate reasoning interfaces, which scale poorly as more reasoning forms are incorporated. This raises a fundamental question: can heterogeneous forms of visual reasoning be unified in a single shared latent space? We investigate this question through HVLR (Heterogeneous Visual Latent Reasoning), which unifies eight distinct reasoning forms within a common autoregressive latent interface. To reconcile heterogeneous supervision, we first render all reasoning representations into a shared pixel space and encode them with the same vision encoder. We then use a two-stage distillation procedure to translate these representations into autoregressive latent states and internalize them into a closed-book model that requires no explicit intermediate representation at inference time. We conduct a systematic analysis with HVLR and uncover three main findings. First, heterogeneous reasoning forms can be effectively unified: combining all eight improves Qwen2.5-VL-7B from 64.66% to 69.97% average accuracy across nine benchmarks. Second, the shared latent space consistently benefits from reasoning diversity. Under a fixed 40K-example budget, increasing the number of reasoning forms from 1 to 3, 5, and 8 improves performance from 66.65% to 67.67%, 68.61%, and 69.97%, respectively, with positive trends across all nine benchmarks. Third, latent reasoning outperforms explicit intermediate decoding while reducing intermediate computation from roughly 177 generated tokens to only 8 latent steps, suggesting that continuous latent representations provide a more expressive and efficient medium for heterogeneous visual reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.