Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models
Abstract
We present a feedforward graph architecture in which heterogeneous frozen large language models serve as computational nodes, communicating through a shared continuous latent space via learned linear projections. Recent work has demonstrated geometric compatibility between independently trained LLM latent spaces, suggesting that cross-architecture activation translation is possible via linear projection. This work extends that finding from static two-model steering to end-to-end trainable multi-node graphs, where projection matrices are optimized jointly via backpropagation through residual stream injection hooks. Three small frozen models (Llama-3.2-1B, Qwen2.5-1.5B, Gemma-2-2B) project into a shared latent space injected into two larger frozen models (Phi-3-mini, Mistral-7B), whose representations feed a cross-attention output node. With only 17.6M trainable parameters against approximately 12B frozen, the architecture achieves 87.3% on ARC-Challenge, 82.8% on OpenBookQA, and 67.2% on MMLU, outperforming the best single constituent model by 11.4, 6.2, and 1.2 points and outperforming majority vote of the same five models by 8.2, 8.0, and 10.4 points — benchmarks where output-level ensembling actively degrades performance. Gradient flow through frozen model boundaries is empirically verified to be tractable, and the output node develops selective routing toward Phi-3-mini without supervision, validated by a 7.8pp performance collapse upon removal.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.