acceptodds
Under review as a conference paper at ICLR 2027

HyperBridge: Effectively Adapting Hypernetworks across LLM Architectures

Abstract

Hypernetworks have recently been explored to internalize long contexts into Large Language Model (LLM) parameters, where a hypernetwork maps documents to parameter updates that enable the adapted LLM to answer relevant queries without repeatedly processing the original context. However, obtaining such a hypernetwork requires substantial data and training compute. Beyond that, the resulting hypernetwork is tightly coupled to its corresponding LLM: its internal modules must match the backbone’s hidden dimensions and number of layers. This limits practical usage under constrained resources since repeating the same training process for each newly introduced LLM is prohibitively expensive. Thus, we ask: Under limited data and compute, can we transfer a trained source hypernetwork across backbones and enable more effective target hypernetwork training? To answer this, we conduct a systematic cross-architecture analysis using representational similarity analysis. We reveal a clear separation in transferability across the hypernetwork: its intermediate representations exhibit substantial shared structure across backbones and can be aligned via a linear mapping. In contrast, its parameter generator remains strongly architecture-specific. Guided by these findings, we propose HyperBridge, a cross-architecture hypernetwork transfer framework that aligns input representations, preserves transferable weights, and re-learns backbone-specific heads. We evaluate HyperBridge across multiple LLM architectures and QA datasets, showing up to 28.5% higher performance than training from scratch and remaining effective across data budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.