ACNA: Decentralized Active-Cell Routing for Sparse and Dynamically Expandable Neural Networks
Abstract
Sparse neural networks activate only a subset of computation, but conventional Mixture-of-Experts (MoE) systems typically rely on a learned centralized router coupled to the expert pool. We propose the Active Cell Neural Architecture (ACNA), in which independently parameterized Expert Cells maintain local activation evidence through a receptor and resting bias, while low-rank Social Intervention provides lightweight inter-cell modulation. A residual transition branch enables function-preserving insertion into pretrained Transformers. Experiments provide a proof of feasibility. A controlled CPU ablation shows lower latency when inactive experts are skipped. In Qwen2.5-Coder-32B, two-cell Python/SQL routing reaches 97.5%; a Math cell can be inserted and later removed without retraining the remaining cells; and mixed-domain prompts place both relevant cells in the Top-2 in 100% of 60 cases. Sequential expansion to eight domains reaches 159/160 (99.38%) in a post-calibration routing evaluation, with 80/80 (100%) accuracy on the original four domains; because this 160-prompt pool overlaps routing-calibration material, we do not label this result blind or held-out. An isolated eight-expert NVIDIA GB10 benchmark yields 1.09× speedup and 8.44% lower latency, while energy/token is 3.32% higher under the current PyTorch execution path. Complementary RTL simulations verify one-hot cell gating, dynamic cell insertion, and a reduction in registered compute activity from 300 dense events to 100 sparse events (66.7%). The RTL result is an activity-level hardware feasibility result, not a direct board-power measurement; together with the GPU result it motivates future cell-aware accelerators with grouped dispatch and fine-grained clock/power gating.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.