DepLatent: Visual Dependency-Guided Latent Multimodal Reasoning
Abstract
Explicit chain-of-thought (CoT) improves multimodal reasoning but incurs substantial inference cost with long and verbose traces. Continuous latent reasoning reduces the number of generated tokens, yet it often underperforms explicit CoT on visually grounded tasks involved with multi-step planning. We first investigate this performance gap through fine-grained error analysis and image-masking interventions, suggesting that latent states may be weakly tied to task-relevant visual evidence and exhibit limited step-wise variation. In this paper, we introduce DepLatent, a dependency-guided latent supervision framework that uses region-wise image masking during explicit CoT reasoning to estimate dependencies between reasoning tokens and visual regions. These estimates provide two complementary training signals: dependency-weighted alignment between latent states and relevant visual features, and step-level reconstruction of CoT segments identified by changes in visual dependency. At inference, the multi-modal model reasons using compact continuous latent tokens, avoiding lengthy CoT generation. Across six multimodal reasoning benchmarks and three vision-language backbones with varied sizes, DepLatent consistently improves reasoning accuracy over explicit CoT and latent reasoning baselines while significantly reducing inference latency. Notably, it achieves an average 6.2% improvement over the original Qwen2.5-VL-7B model and speedups over explicit CoT of 2.47 on the BLINK benchmark.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.