acceptodds
Under review as a conference paper at ICLR 2027

Beyond Feature Fidelity: Recovery and Readout Adaptation in Multimodal Large Language Models

Abstract

Feature recovery targets the geometry of a clean visual representation; multimodal prediction depends on how a readout uses that representation. We examine this distinction in controlled systems connecting a shared visual encoder to frozen language-side models through learned readouts. Across four downstream models and nine benchmarks, recovery has mixed effects under a clean-trained readout. Readout adaptation improves the reported utility in most evaluated model–benchmark combinations, yet adapting directly to raw features is often competitive with or better than adapting after explicit recovery. In a GQA control, additional readout training on recovered clean features raises the score from to , while adaptation using recovered clean and corrupted features reaches . This limits attribution of the entire adaptation gain to distribution matching. Separate fidelity analyses characterize task-specific and task-agnostic recovery settings. Together, the results motivate evaluating the representation–readout pair through distinct measurements of feature fidelity, fixed-readout utility, and adapted-readout utility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.