acceptodds
Under review as a conference paper at ICLR 2027

Wider by How Much? Learning Rates Shape Vision-Language Adapter Comparisons

Abstract

Adapter width in frozen-backbone vision-language models is often chosen from a width sweep trained at one shared learning rate. We show that the measured advantage of wider adapters depends strongly on that rate. Across five Perceiver-adapter widths and frozen language models from 1.5B to 14B parameters, the widest adapter beats the narrowest by 0.7–1.0 nats under a fixed rate at 1.5B–7B; selecting the rate at one width and sharing it removes 69–90% of that gap, and at 14B the narrowest cell varies by 0.57 nats across runs, so no ordering is determined there. Under the transferred rates the widest adapter is best or within 0.033 nats of best and 0.13–0.25 nats ahead of the narrowest. Changing the shared rate also changes Q-Former width comparisons, and the fixed-rate deficit extends to answer likelihood for the four narrower widths at 3B and all five at 14B, and to VQA accuracy after instruction tuning at 3B. A width-adjusted coordinate, the learning rate times the width, predicts held-out validation loss with 34–60% lower error than the rate alone below 14B. A one-anchor rule that rescales the reference rate by the square root of the adapter parameter-count ratio removes the per-width search, and at 14B its rates coincide with the sampled bracket minima at the two widest adapters. Matching the anchor's effective weight decay removes much of the rule's advantage at wide adapters, so the rule transfers the learning rate together with its decay. At 14B a single coordinate leaves width-dependent residuals; clipping settings barely change the narrowest adapter's loss at its rule rate, whereas halving the head count of the widest adapter at fixed parameter count changes its loss by to across two seeds. Connector comparisons should report learning-rate sensitivity or transfer the rate across widths.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.