acceptodds
Under review as a conference paper at ICLR 2027

ViPLE: Visual Inference with Parallel Latent Encoding

Abstract

Explicit visual reasoning generates intermediate text sequentially, increasing decoding cost. We investigate whether useful visual reasoning context can instead be constructed in parallel for a frozen multimodal backbone. We propose VIPLE (Visual Inference with Parallel Latent Encoding), which extends one-pass latent trajectory encoding with a final cross-attention read of fused image-question states. The latent slots attend to these states before conditioning the frozen decoder. Supervised alignment, on-policy distillation, and reinforcement learning train the interface while keeping the backbone frozen. Across six benchmarks, on Qwen3-VL-2B, 4B, and 8B, VIPLE improves macro accuracy by **+(5.12–7.15) pp** over the cheapest measured explicit-reasoning settings while decoding fewer tokens. On GLM-4.1V-9B and Qwen3-VL-32B, the gains are +2.31 pp and +2.50 pp, respectively, at matched decoding costs obtained by interpolation. The cheapest measured explicit-reasoning settings that reach VIPLE's accuracy require **(1.8–5.8)×** as many decoded tokens. Length-preserving replacement controls show the importance of learned latent content, while retraining without the final cross-attention layer reduces accuracy on mathematical tasks. These results demonstrate that a learned parallel interface can improve the accuracy-decode-cost trade-off of frozen multimodal models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.