TraceTilt: Context-Supervised Model Editing and Steering
Abstract
Context changes what a language model does and knows; we study how to retain these changes after the context is removed. We introduce , a framework for context-supervised steering and model editing. We use model editing to construct steering from the predictions of a Context Teacher, the same model given additional context. The retained behavior can itself acquire information: a question-rewriting operation constructed from four reference contexts transfers to new records. For knowledge editing, four reference contexts with example questions initialize reusable task behavior; each new document then supplies its own editing supervision, without access to future evaluation questions. The resulting task and information updates are merged into ordinary weights. On three behavioral tasks, achieves 92.2% mean target-choice accuracy, versus 70.4% for the strongest evaluated baseline. On unseen questions without the document, the edited model recovers 28.6% of the Context Teacher's improvement over Base on BFCL and 26.9% on FictionalQA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.