acceptodds
Under review as a conference paper at ICLR 2027

LayerRefit: Final-Relative Refinement for Affective Multimodal LLMs

Abstract

Encoder hierarchies contain affective cues whose utility varies across modalities and conditions. We introduce LayerRefit to adapt pretrained multimodal large language models (MLLMs) by converting this evidence into residual corrections of their native modality tokens. Depth-specific evidence projections transform differences from the final encoder state; a shared correction projection maps them into token updates, combined with input-dependent weights. Backbone weights and input sequence length remain fixed. Across three affective backbones, source-trained branches improve 22 of 27 cross-benchmark evaluations. Models trained with audio, vision, and text improve 11 of 12 inference-time input configurations on the two primary backbones without retraining, using fewer than 0.6M added parameters. Ablations favor final-relative evidence with shared token correction, and norm-matched interventions show that learned update directions improve recognition. Code: https://anonymous.4open.science/r/LayerRefit-7263/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.