acceptodds
Under review as a conference paper at ICLR 2027

Cross-Model Stitching Transfers Answers, Not Computation

Abstract

Activation stitching describes the technique of grafting one model's hidden state into another's through a learned map. Its motivation is related to universality, under which models may learn shared representations that enable transfer. Prior work demonstrates that features, probes, and steering vectors can survive such transfers. However, what actually moves through the stitch, a question with substantial implications for universality, remains unresolved. We examine this question by grafting a large donor's residual state into a smaller recipient using a map trained with cross-entropy on the gold output label. We spread this evaluation across three same-family model pairs (Gemma-2 9B2B, Qwen2.5 7B0.5B, Llama-3.1 8B3.2 1B) and two cross-family pairs (Qwen2.5 7B and Llama-3.1 8B into Gemma-2 2B). We evaluate across five task families: arithmetic, symbolic binding, state tracking, naturalistic factual recall, and GSM8K. Our experiments find that this task-supervised stitch is a working channel that confers the donor's answer on up to of otherwise unsolvable problems in our primary arithmetic evaluation. Additionally, a battery of controls indicates that conferral is achieved through the transfer of a precomputed answer, and success is dictated primarily by the answer's presence at the donor stitch site. We also find that this presence alone is not sufficient, as it must then be lossily encoded into the recipient. Further ablations indicate that the task-supervised map learns to mitigate this effect by writing the signal redundantly across hundreds of dimensions of the graft, even though the signal itself is low-rank. Finally, while the first answer token transfers exceptionally well from the prefill state, the rest of the answer does not until prior answer tokens have been emitted. Transfer is bottlenecked by answer presence, and the donor must undergo additional serial computation before that presence is sufficient for transfer. Together, these results suggest a boundary in cross-model universality, where models appear sufficiently aligned for precomputed answers to transfer through a stitch, while showing no measurable evidence that the computations producing those answers transfer with them.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.