acceptodds
Under review as a conference paper at ICLR 2027

Story-Gradient Editing Can Misattribute Newly Learned Real Facts

Abstract

Language models learn from factual records and fiction. They need to remember both without confusing which events are real. Epistemic Goggles is an existing method that edits training gradients so that models remember learned material as fictional. We test whether this framing reaches new real facts learned alongside the stories. Matched Qwen3-8B runs share initialization, documents, and update count, and only story updates use the editor. Edited runs retain real awards but report them as fictional even when ordinary training learns their correct status. Synthetic events with varied relations and persons show the same separation between content and attribution. Correct source reports also fail to ensure consistent judgments and applications of the same events. Source supervision improves some answers without making these behaviors agree. The attribution effect also reaches events absent from the edited documents. At half amplitude, both tested synthetic donor updates change correct real-award judgments and retain award content. Three spectrum-preserving random directions retain the original judgments. These results show why selective learning requires more than retaining prior knowledge or reproducing a source answer. Newly learned facts must acquire their own status and support its use across questions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.