GOGGLES: When Output Serialisation Changes the Benefits of Higher Rank in VLMs
Abstract
Vision Language Models (VLMs) can acquire dense visual perception skills by distilling specialized expert targets via Low-Rank Adaptation (LoRA). However, because autoregressive VLMs generate spatial predictions sequentially, a fundamental question remains: how does adapter capacity interact with output serialisation, the way structured targets are formatted into text token sequences? We address this with GOGGLES, a framework that evaluates detachable LoRA adapters on frozen VLM backbones. Our design philosophy isolates output serialisation from information content by holding teacher target resolutions strictly fixed while varying output organisation. We uncover three primary insights: i. Serialisation-Rank Synergy: lossless output restructuring disproportionately benefits low-rank adaptation, allowing low-rank models to approach the fidelity of high-rank counterparts; ii. Domain Specificity: this interaction is highly modality-dependent, holding strongly for continuous depth rasters but vanishing in discrete segmentation, revealing that sequence dynamics depend on spatial target structure; and iii. Fidelity-Efficiency Trade-offs: gains in expert-relative fidelity require additional inference prefills and compute, demonstrating that serialisation design must balance fidelity with deployability. Overall, GOGGLES establishes output serialisation as a critical dimension alongside adapter rank, demonstrating how target sequence structure can alter rank requirements in VLM perception.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.