Construct, Replay, and Reason: Visual-Semantic Latent Block for Multi-modal Reasoning
Abstract
Latent Visual Reasoning (LVR) performs intermediate computation in a visual latent space, allowing reasoning states to carry visual information without verbalizing it. It constructs this space by aligning hidden-state-derived latents with selected visual tokens. However, our analysis identifies two limitations: direct alignment overlooks the distinct representational roles of high-level hidden states and local visual tokens, while using ground-truth RoI visual tokens during training but self-generated latents during inference introduces a substantial training–inference gap. We propose **CR**, a two-pass framework that **C**onstructs and **R**eplays visual-semantic latent block for multi-modal **R**easoning. In Pass-I, a latent head transforms sampled seeds into a multi-token latent block, with auxiliary supervision from RoI visual features and answer-related content. The constructed latent block, serves as the reasoning state and is replayed in Pass-II for answer generation. CR employs the same construct-and-replay procedure during training and inference, thereby reducing their discrepancy. Experiments with 3B and 7B backbones across six multimodal benchmarks demonstrate consistent improvements. At the 7B scale, our method outperforms the strongest baselines by 3.3% and 3.2% on RealWorldQA and OCRBench, respectively, and by 6.4% and 3.5% on BLINK Counting and Spatial Relation. Our project is available at https://anonymous.4open.science/r/CR2.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.