acceptodds
Under review as a conference paper at ICLR 2027

Bridging the Representation–Behavior Gap in Knowledge Editing with Calibrated Activation Steering

Abstract

Knowledge editing aims to update specific knowledge in large language models while preserving unrelated behavior. Recent distribution-based approaches edit knowledge by steering hidden representations toward distributions associated with the target knowledge. However, it remains unclear whether better representation alignment necessarily leads to the desired output behavior. Through systematic layer-wise analysis, we identify a **representation–behavior gap** with two distinct failure modes: **behavioral under-realization**, where well-aligned hidden representations fail to sufficiently express the target knowledge in output behavior, and **behavioral overshooting**, where stronger interventions continue to improve edit success but push the output behavior beyond that induced by explicit conditioning on the target knowledge. To address these failure modes, we propose **B**ehavior-aware **E**diting via **C**alibrated **A**ctivation**S**teering (**BECAS**). BECAS uses the model's output distribution when explicitly conditioned on the target knowledge as a behavioral reference and adjusts the strength of activation steering accordingly. Experiments on knowledge editing benchmarks show that BECAS improves generalization and locality while maintaining high editing efficacy. These results demonstrate that output behavior provides a useful reference for controlling representation-based knowledge editing.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.