BEYOND 4D FFNS: HESSIAN POTENTIAL MLP FOR MEMORY AND LANGUAGE MODELING
Abstract
Position-wise feed-forward networks (FFNs) dominate Transformer cost, yet remain standard two-layer MLPs expanding to width before linear projection. We introduce the Hessian Feed-Forward Network (H-FFN), , expressing the sublayer as a sum of rank-one directional units derived from a scalar potential. For associative retrieval, a coaxial Memory-Track employs multi-scale contextual gating; for language modeling, Decoupled H-FFN separates content, routing, and emission. Under matched training budgets, Memory-Track reaches Needle-in-a-Haystack accuracy versus for a matched MLP, and Decoupled H-FFN achieves – perplexity gains across causal language modeling benchmarks ( on WikiText-103; on FineWeb-Edu). In transfer, analytic GELU matching improves GPT-2 fine-tuning () and accelerates ViT-B/16 convergence ( top-1 with fewer steps).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.