The Normalization Landscape in Continual Learning: Resolving Moving-Average Breakdown and Subspace Interference for Buffer-Free Edge Intelligence
Abstract
Modular parameter isolation structurally eliminates catastrophic forgetting in Task-Incremental Learning: f(x; theta^(t+1), M_k, mu_k, sigma_k^2) = f(x; theta^(k), M_k, mu_k, sigma_k^2). Yet modular continual learning collapsed on moving-average networks due to the Normalization Paradox: in BatchNorm architectures, pruning channels to isolate sub-circuits causes dense running moments to mismatch sparse activations, shifting pre-activations negative and clamping up to 98% of post-ReLU signals to zero. In autonomous Class-IL, modular networks lack test-time routing without raw exemplar buffers (|M| = 0), where off-chip DRAM fetches dissipate > 200x higher energy than on-chip MACs. Sparse Mechanistic Routing (SMR) resolves both challenges for resource-constrained edge systems: (1) O(1) BatchNorm Recalibration realigns running statistics over K = 20 micro-batches in under 0.7 s without backpropagation, reviving dead sub-networks (+57.76% on MobileNetV3) and restoring Neural Collapse ETF geometry; (2) Sensory-Decision Partitioning decouples shared visual primitives from isolated decision circuits; (3) Autonomous Test-Time Routing executes an O(1) Fast-Gate (10.10 ms, 9.21x speedup) with conformal triage (>= 95.0% coverage via q_alpha; 81.4%-89.2% under |C(x)| <= 2). Crucially, Orthogonal Subspace Multiplexing (SMR-OSM) multiplexes orthogonal weight bases without pruning (rho_in = 1.0), proving that orthogonal multiplexing completely avoids the Normalization Paradox (Delta mu = 0) by design with 0% parameter growth, reaching 61.40% Task-IL and 38.80% Class-IL on Split CIFAR-100. With Dynamic Modular Expansion (D-SMR, +18.7% width, +96 channels), temperature-calibrated conformal triage elevates Class-IL to 56.20%, decisively surpassing DER++'s 55.10% replay baseline without storing a single exemplar (|M| = 0), while retaining 0.00% forgetting when frozen. Microsecond profiling on physical Blackwell GPU silicon confirms an empirical 2.01x step energy reduction (2085.74 vs 4183.26 mJ) and 2.01x throughput speedup, while a compile-time static pre-allocated envelope guarantees strictly 0.0000 KB dynamic runtime heap allocation, fitting T = 50 metadata (1.44 MB) within edge SRAM (< 2 MB) where 28nm CMOS microarchitectural modeling projects up to a 13.53x cumulative energy reduction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.