Rethinking the Mamba Gate Projection
Abstract
We study whether Mamba's independent gate projection is necessary, or whether its routing signal can instead be derived from a representation the block has already computed. We propose Content-Derived Gating (CDG): a low-rank map applied to Mamba's content projection in place of the full gate. CDG substantially reduces gate parameters while preserving performance below flagship (130M-parameter) scale, where it trains 19% faster and uses 23% less peak memory than the full gate but incurs a persistent 2.6% relative bpb gap. We then show this behavior is not explained by expressivity: CDG and a matched content-blind low-rank gate parameterize exactly the same set of reachable gate functions whenever the content projection has full rank, as it does throughout this study (Proposition prop:reachability). Yet the two behave differently in practice, and the direction reverses with scale: comparable or suggestively better for content-derivation at small scale, but at 130M the content-blind gate improves bpb by over CDG despite fewer parameters, closing more than half its gap to the full gate. Any derived-routing scheme's gains should therefore be checked against a content-blind low-rank baseline of matched size before being credited to the derivation, since that credit can reverse with scale.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.