acceptodds
Under review as a conference paper at ICLR 2027

Can Topological Decoding Make VLMs as Reliable as PageRank Made Web Retrieval?

Abstract

Vision-Language Models (VLMs) frequently suffer from generative semantic drift, where a heavy reliance on frozen parametric weights overrides discrete visual evidence, causing compounding hallucinations. Existing decoding interventions attempt to mitigate this but often fracture linguistic fluency or incur prohibitive latency. To resolve this factuality-fluency-latency trade-off, we introduce Topological Decoding, an inference-time architecture that fuses discrete Visual Grounding Graphs (VGG) with the VLM's logits via a PageRank-inspired algorithm. To accurately localize the VGG and achieve high spatial precision with minimal latency, we employ Group Relative Policy Optimization (GRPO) for VLM object recognition. During inference, a Controllable Faithfulness Dial () is introduced to provide users with dynamic flexibility, interpolating between factual grounding and linguistic fluency. Adding only 0.3 seconds of latency, Extensive evaluation on 1.2 million Visual Genome queries demonstrates our approach consistently outperforms state-of-the-art interventions in spatial faithfulness, achieving 72.8% compared to VCD (52.1%), OPERA (56.4%), and Gradient-based decoding (59.2%). Furthermore, our RL-adapted policy achieves 94.1% [email protected] on RefCOCO while fully preserving native multimodal capabilities on MT-Bench and MMMU.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.