When Transformers Copy Instead of Compose: Circuit Competition under Shared Semantics
Abstract
Foundation models are often given the same semantic event realized in several modes at once: an image and its description, a sentence and its translation, a specification of a program and its implementation. The semantics are shared across views while the realization is not, so a model can form two competing circuits: one that computes the content in the target view and one that copies it from the source view. We study this phenomenon in detail with synthetic data generated from probabilistic context-free grammars (PCFGs). Specifically, we hold view difficulty fixed and dial the levels of shared structure: a controlled hierarchical setting where two views can share semantic structure down to a precisely controlled level , with exact Bayes-optimal ceilings determining how much information is available from each view at each level. Probing the residual stream at every layer, we find the following: First, fusion between the views happens at the deepest shared latent level (the fusion boundary), the level where the shared information first becomes sufficient for the target view. Second, the availability of a copying route weakens composition at semantic levels above : once the model can copy the shared content from the source view, its further composition in the target view weakens. Third, this weakening is reversible: removing the target's access to the source, or otherwise reducing the copying route's reliability, restores the missing computation, which coexists with the copying route rather than being lost. Finally, we observe a consistent pattern in LLaVA-1.5-7B, where blocking attention to a conflicting caption restores image-based answers, and discuss implications for fusion in multimodal models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.