When Can One Transformer Use Another Transformer's Hidden State?
Abstract
When can separately optimized transformers directly use one another's hidden states without learned alignment? We study raw interoperability: whether a frozen recipient can causally use an independently learned donor state under the identity map. In controlled retrieval tasks, independently optimized models sharing externally specified coordinate scaffolds exhibit robust cross-model transplantation, including on a held-out two-hop compositional task. Decoder-subspace analysis shows that decoder-visible state carries substantial but incomplete transferable signal, while scaffold interventions identify shared output coordinates as the strongest tested anchor; recipient-specific decoder-null structure remains consequential. When output coordinates differ, identity-map transfer nearly disappears, whereas applying the exact known coordinate correspondence restores nearly all of the lost advantage as mapped compatibility. In a complementary boundary test, independently pretrained Pythia-70M seeds show no reliable cross-seed identity-map transfer at the tested interfaces, including a site subsequently verified in a post hoc diagnostic to be causally informative within individual models. These results identify raw cross-model interoperability as a scaffold-dependent interface property: it can emerge across independently optimized models, but depends on shared coordinate conventions and should not be assumed without direct testing in independently pretrained networks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.