acceptodds
Under review as a conference paper at ICLR 2027

Generalizable Model Editing Could Be More Than Direct Generalization

Abstract

A central desideratum in model editing is generalization to complex tasks, which prior work formalizes under a *direct generalization* paradigm: treating and evaluating LLM behavior as an input-output mapping and expecting a factual update to alter outputs on reasoning tasks. Yet despite substantial empirical effort, the field consensus remains that such generalization is highly challenging and elusive. However, when we deploy the latest models in an in-the-wild chatbot-user setting, we surprisingly find that even classical editing methods exhibit remarkably fluent generalization behavior. This motivates us to revisit the obstacles to generalization and the limitations of the prevailing paradigm. We find that direct generalization is fundamentally constrained both theoretically and practically: theoretically, it suffers from an identifiability issue; and practically, existing approaches attempting to circumvent this are further hindered by the ambiguity of their constructed biases and by vulnerability to being overridden by other factors such as optimization dynamics. We instead identify *compositional generalization*, empowered by the increasingly strong reasoning and composition abilities of modern LLMs, as a theoretically grounded and empirically validated alternative. We validate this phenomenon across multiple models and editing methods, showing that such generalization behavior is widespread. Overall, we argue that generalizable model editing could be more than direct generalization, and call for a paradigm shift toward more generalizable and practically deployable knowledge update techniques.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.