APERTURE: Few-Shot Refocusing of Visual Evidence in Vision–Language Models
Abstract
In a new visual domain, a vision–language model may need to make many predictions using only a few labeled images. This situation arises, for example, when identifying tumor tissue at a new hospital or assessing building damage after a new disaster. We introduce APERTURE, a small, reusable module that adjusts image-token key/value states while keeping the backbone frozen. Gradients of the correct-answer loss on labeled support images identify sensitive directions in these states. The module learns a bounded transformation within those directions and applies it to each new image’s activations, without demonstrations or further optimization. Across four BRIGHT events and three support sets per event, APERTURE increases Qwen2.5-VL’s mean macro-F1 from 22.55% to 31.33% using only 1,024 learned coefficients, compared with 31.46% for LoRA-1E. On CAMELYON17–H2, it achieves lower mean raw candidate negative log-likelihood than LoRA-1E and LoRA-4E across the three evaluated Qwen model variants. In the measured Qwen2.5-VL configuration, its saved adaptation state, including the fixed bases, occupies approximately 1/291 of the storage required by rank-16 LoRA. These results show that APERTURE can achieve performance close to LoRA-1E on BRIGHT with substantially less adaptation storage, while its benefits vary across tasks and models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.