ELVIS: Evidence-Grounded Latent Visual States for Multi-Step Reasoning
Abstract
Multi-step multimodal reasoning requires sustained access to visual evidence, yet recurrent methods, constructing latent visual states sequentially, show weak sensitivity to decisive evidence and provide limited support for subsequent reasoning in our diagnostics. We introduce ELVIS, a framework that constructs evidence-grounded latent visual states in parallel and invokes new states as reasoning progresses. Learnable queries interact bidirectionally within each state while preserving causality across reasoning steps. To supervise state content, we construct AVR-112K, an Action–Visual–Readout corpus whose explicit observations guide teacher-based evidence selection. Evidence alignment and visual anchoring train latent states to capture the selected evidence and attend to the original image. An Invocation Progressive Reward encourages additional invocations when they yield higher empirical success than smaller observed invocation counts. Experiments demonstrate strong performance across visual perception and spatial reasoning tasks, including a 22.04% gain over Qwen2.5-VL-7B on Zebra Maze. At 16 latent positions and matched invocation counts, ELVIS achieves an 11.65–17.40× speedup in latent-state computation over the fastest applicable recurrent baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.