APERTURE: What to measure, how much to measure, and where to allocate in multimodal low-rank adaptation
Abstract
The same multimodal input may require different evidence depending on the question, yet conventional low-rank adaptation (LoRA) exposes the same fixed set of measurements regardless of the query. Under a limited adaptation budget, this fixed allocation can miss question-relevant evidence, provide insufficient within-modality support, and prevent capacity from shifting across modalities as their relative importance changes. We formulate this challenge as question-conditioned measurement allocation: jointly deciding what to measure, how much to measure, and where to allocate a shared average measurement budget. We introduce Aperture, a framework that operates on a fixed bank of reference and gradient-informed candidate directions. Question-conditioned rotations adapt measurement directions, gates control the amount of exposed support within each modality, and a joint support regularizer redistributes the shared budget across modalities. In experiments with LLaMA-2-7B, Aperture outperforms the question-blind Static-Alloc control by 4.18 percentage points on the official MUSIC-AVQA split and 2.23 points in AVE segment accuracy, at matched mean support and trainable parameter count. On a video-disjoint MUSIC-AVQA split, it retains a 2.09-point gain. These results demonstrate the benefit of making measurement allocation question-dependent under a fixed adaptation budget. Code is available at https://anonymous.4open.science/status/Aperture-F0C0.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.