PolyGate: Continuous Power-Product Gating for Deep Neural Networks
Abstract
Standard feedforward and convolutional neural networks parameterized with piecewise-linear activations approximate smooth polynomial and multiplicative interactions through piecewise-linear tessellations, whereas multiplicative architectures can represent power-law and cross-variable interactions directly. However, historical product-unit formulations suffered from unstable optimization, logarithmic singularities near the origin, and undefined real-valued fractional powers on negative inputs. In this work, we propose PolyGate, an architectural primitive that augments affine coordinate bases with learnable continuous power-product gating. To ensure well-defined real-valued power operations and exclude zero-boundary logarithmic singularities from the operating domain, we introduce Strict-Positivity Batch Normalization (SP-BN) and explicit positive-domain handling. We further introduce dimension-normalized multiplicative gating to control activation variance as feature dimension grows. Theoretically, we characterize exact representation of univariate and multivariate generalized monomials, gradient regularity on strictly positive bounded domains, near-identity initialization with asymptotic recovery of the affine baseline, and dimension-scaling laws of multiplicative gates. Empirically, PolyGate reduces test MSE by on Feynman relativistic mechanics, achieves higher mean/median out-of-distribution across the complete 10-equation AI-Feynman suite, and is evaluated across convolutional backbones on Tiny-ImageNet-200: delivering to gains on edge MobileNetV4 with modest parameter overhead ( to ), and to gains on ConvNeXt with dual depthwise-pointwise gating.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.