Deformable Vision Query Connector: Letting LLMs Query Vision Information through Position-Grounded Lookup
Abstract
Large Vision-Language Models (LVLMs) represent images as long sequences of visual tokens. However, preserving fine-grained visual content under tight to- ken budgets remains challenging: spatial aggregation often lacks adaptivity, while query-based resampling can collapse without extensive training and miss impor- tant regions. We propose the *Deformable Vision Query (DVQ) Connector*, a position-grounded and spatially adaptive interface embedded directly within LLM layers. DVQ associates each visual query with a reference point and uses de- formable cross-attention to refine its spatial focus. Uniformly sampled reference points encourage image coverage, while learned refinements allow adaptive con- centration on salient regions, overcoming limitations of cross-attention connectors through this position-grounding. DVQ also natively encodes positional informa- tion within its queries, enabling differentiable bounding box encodings. Exper- iments on small-scale per-task fine-tuning and general-purpose vision-language modeling demonstrate that DVQ is competitive with strong vision connectors like DeepStack, while DVQ intrinsically enables token compression, making it partic- ularly effective under reduced vision token budgets and in high-resolution settings. Ablations further demonstrate DVQ’s robustness and flexibility. We provide our source code at https://anonymous.4open.science/r/dvq.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.