LatentWeave: Learning High-Value Thought for Efficient Visual Reasoning
Abstract
Scaling reasoning increasingly relies on longer chains of thought. Because natural language expresses both reasoning progress and elaboration of existing thought, scaling explicit reasoning does not necessarily constitute more intelligence. We hypothesize that reasoning capability depends on high-value thought and its coherent organization. Motivated by this hypothesis, we introduce \method, combining rubric-guided reinforcement learning to retain high-value explicit reasoning with interleaved latent updates to maintain reasoning continuity. Experiments on six benchmarks covering hard multimodal, vision-dependent, and long-chain reasoning establish an empirical efficiency–accuracy Pareto frontier. Relative to the Qwen3-VL-8B-Thinking baseline, \method achieves reasoning efficiency with 1.86% relative accuracy loss, and efficiency and 3.84% loss. This Pareto frontier supports flexible model selection on varied reasoning budget and performance requirements.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.