acceptodds
Under review as a conference paper at ICLR 2027

Train Wide, Act Narrow: The Learning Geometry of Visual Adaptation

Abstract

In video-language adaptation, wide visual adapters learn corrections concentrated almost entirely along one output direction, while their hidden features and readout weights remain broad. Across the compression methods we study, retaining that direction preserves nearly all of the average benchmark gain, while removing it returns accuracy close to the unadapted model. We call this pattern *Train Wide, Act Narrow*. This concentration emerges through learned hidden-readout alignment, with the readout preferentially amplifying a dominant mode of hidden variation. A one-hidden-unit adapter constructed from a trained Wide adapter reaches development negative log-likelihood (NLL) of 1.200, while the best task-trained counterpart we test reaches 1.403. We then compare two readouts that compute exactly the same one-axis function from the same hidden features. The unrestricted readout can produce transverse changes using coefficient patterns unavailable to the one-axis readout. Even the best single pattern misses 17.3 to 32.1% of the transverse gradient energy accessible to the unrestricted readout. After continuation from the same one-axis function, projecting the resulting unrestricted correction onto its dominant direction raises NLL despite retaining more than 99.5% of its raw energy. A nearly one-dimensional correction can require more freedom to learn and more than one direction to preserve its behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.