acceptodds
Under review as a conference paper at ICLR 2027

R-MLLM: Dynamic Model Editing for Robust Multimodal Large Language Model under Limited Data

Abstract

In real-world scenarios, unknown visual corruptions may continue to emerge and degrade model performance. Existing robust Multimodal Large Language Models (MLLMs) are typically optimized for predefined corruptions and thus generalize poorly to unknown corruptions. Although adapting these models to new corruptions can improve their performance on newly emerging corruptions, it often leads to catastrophic forgetting of previously acquired robustness knowledge. To address these challenges, we formulate the visual robustness task as a continual learning problem and propose a model editing framework for data-efficient continual adaptation to visual corruptions. Specifically, we fine-tune the visual encoder with an additional robustness loss to obtain corrupted key-value pairs. We then propose a dynamic key-value pair construction strategy that adaptively weights samples according to their quality, enabling effective model updates with limited data. To preserve pretrained knowledge, we further introduce a principal component constraint that regularizes parameter updates during adaptation and mitigates catastrophic forgetting. Finally, we introduce VCBench, a benchmark covering vision questions under 17 common real-world visual corruptions, to evaluate the robustness and continual adaptation ability of MLLMs. Extensive experiments across multiple benchmarks show that our method improves robustness over strong baselines, achieving up to a 7.6% average accuracy gain on VCBench.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.