Treating Alignment as an In-Context Approximation Problem
Abstract
Alignment changes model weights, yet much of what steers model behavior at test time happens in context. Their immediate relationship is so far little understood. We study when context-induced distribution shifts can reproduce those induced by parameter updates. We first show that, in linear attention, a fixed KV prefix exactly reproduces a low-rank bilinear update, and characterize where this fails, including updates whose effects depend on task content. This motivates task-dependent context, where relevant clauses are retrieved from an approved constitution. We identify three conditions for close approximation: the retrieved clauses fit the update, the frozen model realizes their declared effects, and the clauses retain those effects when combined. Controlled experiments agree with this picture: across LoRA updates, retrieving the relevant clause recovers the trained rule in frozen Qwen3 models and brings them much closer to their aligned counterparts, while presenting the full constitution can introduce errors from irrelevant and competing rules. Finally, we identify two fundamental limits of in-context alignment. Exact bounded clause selection is NP-complete in the worst case, and some training-induced shifts lie in directions that no deterministic context can follow.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.