DO STATE SPACE MODELS NEED SEPARATE B AND C PROJECTIONS? REVISITING B-C TYING FROM S4D TO MAMBA
Abstract
We study projection redundancy: whether the distinct linear roles a sequence model uses to read and write information can be collapsed onto shared parameters without losing capability. Diagonal state space models (SSMs) such as S4D, S5, and Mamba read and write their hidden state through two matrices, B (input) and C (output); we test tying C = B across four case studies, multi-seeded throughout except where flagged single-seed. For a static SSM (S4D), tying is a clean free lunch — no measurable cost at any depth or task tested, including a byte-level language model — cutting the SSM layer’s parameters by a third. For a selective SSM (Mamba), tying carries a real ∼14% capacity cost, traced to realvalued B,C forcing a tied mode’s contribution to be non-negative (unlike S4D’s complex-valued state); a fixed sign mask fixes this at zero extra parameters, recovering ∼86% of the gap, and a parameter-matched recursive check confirms the fix is a genuine effect rather than fewer parameters in disguise, beating a matchedbudget free model decisively on a real-corpus language model. A plain RNN with a genuine nonlinearity inside the recurrence again costs nothing reproducible, and a systematic study of attention found tying two of three Q,K, V projections costs < 1% relative accuracy (Kayyam et al., 2026). Post hoc, converting an alreadytrained free model to tied is exact for S4D, damaging but fine-tune-recoverable for Mamba, and on a real language model, over 80% recoverable (single seed) using only the frozen model as a synthetic teacher, without the original training data. Together, these four case studies show projection redundancy is architecturedependent in its mechanics, but when tying costs something, the cause is often diagnosable and cheaply fixable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.