acceptodds
Under review as a conference paper at ICLR 2027

Reinforcing Large Vision Language Models for Multi-Hop Reasoning in Lifelong Knowledge Editing

Abstract

Lifelong multimodal knowledge editing aims to correct outdated or incorrect knowledge in large vision language models (LVLMs) without retraining. Recent studies have significantly improved editing quality in terms of locality, generality, and reliability. However, the ability of post-edit models to perform multi-hop question answering over sequentially injected facts remains critically overlooked. This challenge stems from the confinement of existing supervision signals to the edited fact itself, with no dedicated mechanism for coherent multi-hop reasoning chains. To bridge this gap, we propose **MoE-GRPO**, a reinforcement learning framework built on GRPO for lifelong multimodal knowledge editing over MoE-based editors. Using two-stage singular value perturbation via full SVD decomposition and carefully designed outcome-level rewards, MoE-GRPO provides direct supervision over multi-hop reasoning, bypassing the need to explicitly model intermediate parameter disruption. Extensive experiments on BLIP2-OPT, LLaVA-v1.5, and MiniGPT4 demonstrate that MoE-GRPO achieves substantial relative gains in multi-hop reasoning, while accepting a small, deliberate trade-off on single-hop editing metrics (Reliability, Locality, Generality) that is intrinsic to the multi-hop reasoning objective. Ablation studies reveal practical trade-offs between editing quality and reasoning ability induced by key design choices. Furthermore, MoE-GRPO generalizes consistently across different MoE-based editors, yielding robust reasoning improvements.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.