Shared Memory, Two Routes: Late Residual Identity Routing in Llama 3
Abstract
We test competing predictions of clinical inter-identity memory isolation versus practice-based shared access - stated in advance in our analysis code - in base Meta-Llama-3-8B using shared-context retrieval, layer-wise residual stream alignment, and layer-31 causal interventions. Under shared context, the Host retrieves a secret key placed exclusively in the Tulpa's turn with 0.661 joint probability versus a 0.0049 no-fact prior, confirming shared memory despite explicit non-sharing directives. Layer-31 representational alignment shows a cosine gap on a short prompt pair (0.355 Host/Tulpa vs. 0.500 Speaker A/B), but twelve matched items narrow means to 0.418 vs. 0.473 (unlabeled: 0.393), and a long mathematics passage closes the gap entirely (0.724 vs. 0.723). Finally, layer-31 zero-ablation recovers only what unembedding readout guarantees, while bidirectional residual swapping yields text degeneration rather than persona transfer. Together, these results confirm that memory is unpartitioned across dual-identity prompts, while showing that persona framing operates merely as a weak terminal-layer constraint on a single residual stream, with no evidence of an internal modular routing mechanism.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.