Think in Rows and Columns: Decoupled Latent Visual Reasoning for Multimodal Tables
Abstract
Multimodal table reasoning requires Large Vision Language Models to jointly recognize fine-grained cell content and reason over structural relations across rows, columns, and headers. Existing methods either generate answers directly or externalize intermediate computation through textual traces, programs, and image operations. While latent visual reasoning offers a more efficient alternative, current methods lack explicit organization according to table structure. To this end, we propose **Decoupled Latent Visual Reasoning (DLVR)**, which organizes continuous latent computation into progressive row, column, and subtable stages while separating structural organization from semantic content. DLVR is trained through decoupled latent supervised fine-tuning followed by **Latent Intervention GRPO**, which employs a **Latent Contribution Reward** to promote the predictive contribution of generated latent components. To supervise this process, we further construct **TableLVR-41K**, a corpus with structure-guided reasoning routes and stage-specific visual and content annotations. Experiments with Qwen2.5-VL and Qwen3-VL across seven benchmarks demonstrate consistent improvements over generic latent reasoning, and explicit table reasoning methods. Ablation studies further validate the complementary contributions of structure-content decoupling, latent supervision objectives, and intervention-based optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.