acceptodds
Under review as a conference paper at ICLR 2027

HASH-EMBEDDING MEMORY INJECTION: CROSS SCALE TRANSFERABLE DIFFERENTIAL MEMORY FOR LARGE LANGUAGE MODELS

Abstract

Fine-tuning is expensive and adapters such as LoRA are tied to one base model, so domain knowledge is voided at every backbone upgrade. We propose Hash-Embedding Memory Injection (HEMI), which stores knowledge in a hash -gram table trained once on a source backbone and reused across backbones of the same family and pre-training lineage through each backbone's anchor basis. Three contributions: (1) a critical-frequency theory giving a closed-form and the table-size bound ; pruning at costs under held-out loss over configurations in domains; (2) differential injection of only the residual between stored memory and the hidden state, under evidential gating and RMS balancing—ablations show the gate and the adapter decide whether a transferred table helps or hurts (removing the adapter hurts in all three settings; removing the gate does so in two and leaves of the gain in the third), while dropping differencing or RMS costs – nats; (3) anchor rendering via thin QR factorization, plus a migration stability theorem bounding direction error by three additive sources: finite-frequency noise, Stage-2 drift, and anchor-subspace mismatch. On DeepSeek-R1-Distill-Qwen, a table trained on B improves the B target in four domains by to nats with target-side parameters ( of LoRA at ), on top of a M-parameter table trained once on the source; the gain does not survive on the B target ( on PubMed, on Math; against on B under a second protocol), where a zero-table control gains more than the trained table ( on PubMed, on Math), so the content of the table contributes nothing there.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.