S2E: Dense Functional Differentiation in Feed-Forward Networks via Prototype-Guided Subspace Modulation
Abstract
Mixture-of-Experts (MoE) architectures improve model capacity via specialized experts and input-dependent routing, but incur parameter duplication, routing complexity, and irregular computation, motivating a fundamental question: can functional differentiation emerge within a single dense feed-forward network (FFN) without independent expert parameters, token dispatch, or sparse activation? We propose Semantic Subspace Expert (S2E), a dense FFN that partitions a shared hidden representation into contiguous subspaces and achieves dense functional differentiation via prototype-guided adaptive modulation. Modulation weights are computed from two complementary signals: global cosine similarity for subspace-level alignment and periodic coordinate-level alignment for fine-grained feature correspondence. Experiments across ImageNet-1K classification, out-of-distribution benchmarks, and COCO detection/segmentation show S2E consistently improves the accuracy–efficiency trade-off with fewer parameters and FLOPs. On DeiT-T, S2E achieves 74.9% Top-1 accuracy with 3.9M parameters and 0.91G FLOPs, outperforming the FFN-4× baseline by 2.7 points with lower costs. Controlled ablations validate the contributions of contiguous subspace organization and dual-signal modulation, confirming shared projection yields higher accuracy and greater subspace diversity than independent per-subspace projections. Further analyses verify non-redundant, category-dependent functional preferences across subspaces while they remain fully active.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.