BEYOND ATTENTION-GUIDED WARPING: LEARNING QUESTION-GUIDED VISUAL REALLOCATION
Abstract
Multimodal large language models (MLLMs)operate under a limited visual-token budget, which can undersample question-relevant details while allocating sub stantial capacity to background regions. Cropping and active visual search can recover local detail, but may remove contextual evidence or require repeated model evaluations. Moreover, intermediate-layer attention does not reliably pre dict which finite image deformation will improve the decoded answer. We propose QGVR (question-guided visual reallocation), a two-stage proposal-and-selection framework that redistributes visual sampling density without changing the frozen MLLM or its token interface. In Stage I, a proposal model predicts a question conditioned, full-frame sampling warp from frozen native-image ViT features, using only image–question–answer supervision without boxes, masks, or region labels. Stage II freezes the proposal model, constructs a candidate set com prising the native observation, the base warp, and eight fixed edits, and uses offline answer scores from the frozen MLLM to supervise native-relative rank ing and fallback decisions. On InternVL, inference requires at most one addi tional ViT encoding of the selected warped image and a single answer-decoding pass. QGVR achieves the highest reported scores on all five InternVL bench marks and four of five LLaVA benchmarks in our comparison. Across four an notated evaluation subsets, QGVR directs the model’s post-decoding visual at tention to target regions more accurately, improving target-hit accuracy by 4.00– 12.18 percentage points over image warping guided by intermediate-layer atten tion maps. Anonymous project materials and a code framework are available at https://anonymous.4open.science/w/qgvr-framework-03F0/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.