Make: Distributionally Robust Knowledge Editing under Multi-Agent Context Shift
Abstract
Knowledge editing for large language models is usually developed and evaluated in isolated settings. In multi-agent systems, role instructions and peer messages change the retrieval context, and thus edits that succeed on clean prompts can fail during interaction. Editing every agent is one fix, but its cost grows with the number of roles. We instead study a setting in which only one designated agent is edited, and ask whether this single intervention can still improve system-level behavior. We identify the main failure mode as interaction-induced representation shift: multi-agent context changes the internal representation used to retrieve the edited fact and can weaken a standard edit. We propose MAKE (Multi-Agent Knowledge Editing), a two-stage framework for multi-agent knowledge editing. Adversarial Context Mining collects deployment-relevant contexts, and Wasserstein-robust multi-key editing enforces the target association on clean and contextual keys while minimizing a tractable upper bound on worst-case off-support activation. We also derive an off-support sensitivity bound for unseen contextual deviations. Experiments on two LLM families and three representative multi-agent architectures show large drops from isolated edit success to final system success for standard editors. MAKE improves multi-agent edit success and conflict robustness while maintaining competitive generalization and preservation under comparable computational cost. It also remains stable under sequential edits, achieving limited degradation of general-purpose capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.