acceptodds
Under review as a conference paper at ICLR 2027

APERTURE: Few-Shot Refocusing of Visual Evidence in Vision–Language Models

Abstract

In a new visual domain, a vision–language model may need to make many predictions using only a few labeled images. This situation arises, for example, when identifying tumor tissue at a new hospital or assessing building damage after a new disaster. We introduce APERTURE, a small, reusable module that adjusts image-token key/value states while keeping the backbone frozen. Gradients of the correct-answer loss on labeled support images identify sensitive directions in these states. The module learns a bounded transformation within those directions and applies it to each new image’s activations, without demonstrations or further optimization. Across four BRIGHT events and three support sets per event, APERTURE increases Qwen2.5-VL’s mean macro-F1 from 22.55% to 31.33% using only 1,024 learned coefficients, compared with 31.46% for LoRA-1E. On CAMELYON17–H2, it achieves lower mean raw candidate negative log-likelihood than LoRA-1E and LoRA-4E across the three evaluated Qwen model variants. In the measured Qwen2.5-VL configuration, its saved adaptation state, including the fixed bases, occupies approximately 1/291 of the storage required by rank-16 LoRA. These results show that APERTURE can achieve performance close to LoRA-1E on BRIGHT with substantially less adaptation storage, while its benefits vary across tasks and models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.