MOEMORY: LONG-TERM USER MEMORY AS A ROUTING POLICY ON FROZEN MIXTURE-OF-EXPERTS LANGUAGE MODELS
Abstract
Long-term language-model memory is commonly implemented by injecting retrieved tokens or updating parameters: injection supplies what to say. We study an orthogonal control mechanism: selection, which governs when to say it by storing user state as a bias over the routing logits of a frozen Mixture-of-Experts model. MoEMory estimates a compact per-user route state from history, then applies either an always-on static bias or a prompt-conditioned graph readout without replaying that history. The bias enters before native scoring and top-k, changing expert selection and weights while expert parameters remain frozen. We derive the geometry of this intervention—from native route containment and top-k stability to exact base recovery on the zero-bias branch—and measure its semantic behavior directly. In a cross-backbone evaluation, static routing raises domain disposition by +0.12 on GLM-4.5-Air, with +0.31 lift for matched memory– query domains and −0.07 for mismatches. Archived 30B comparisons show lower off-topic detector activation for routing at the tested strengths. A controlled coloading sweep separates static from prompt-conditioned behavior, and a timed 30B study measures compact persistent state without history-token replay. Together, these results establish expert routing as a compact control memory for model-realizable dispositions, complementary to content carriers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.