ViPLE: Visual Inference with Parallel Latent Encoding
Abstract
Explicit visual reasoning generates intermediate text sequentially, increasing decoding cost. We investigate whether useful visual reasoning context can instead be constructed in parallel for a frozen multimodal backbone. We propose VIPLE (Visual Inference with Parallel Latent Encoding), which extends one-pass latent trajectory encoding with a final cross-attention read of fused image-question states. The latent slots attend to these states before conditioning the frozen decoder. Supervised alignment, on-policy distillation, and reinforcement learning train the interface while keeping the backbone frozen. Across six benchmarks, on Qwen3-VL-2B, 4B, and 8B, VIPLE improves macro accuracy by **+(5.12–7.15) pp** over the cheapest measured explicit-reasoning settings while decoding fewer tokens. On GLM-4.1V-9B and Qwen3-VL-32B, the gains are +2.31 pp and +2.50 pp, respectively, at matched decoding costs obtained by interpolation. The cheapest measured explicit-reasoning settings that reach VIPLE's accuracy require **(1.8–5.8)×** as many decoded tokens. Length-preserving replacement controls show the importance of learned latent content, while retraining without the final cross-attention layer reduces accuracy on mathematical tasks. These results demonstrate that a learned parallel interface can improve the accuracy-decode-cost trade-off of frozen multimodal models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.