acceptodds
Under review as a conference paper at ICLR 2027

MLP Cache: A Post-Hoc Memory–Compute Trade-off for Transformer Inference

Abstract

MLPs impose substantial inference cost in Transformers by applying a full, high-capacity transformation to every token. Deployed workloads can revisit recurring local MLP input regions, there's opportunity for compute reuse beyond exact-match caching. We introduce MLP Cache, a post-hoc cache that stores compact local predictors of a frozen MLP without retraining the backbone. The central challenge is to allocate limited memory across region scale and predictor capacity: broader regions cover more inputs but may require larger predictors, finer regions are easier to approximate but require more entries. MLP Cache constructs a coarse-to-fine hierarchy using shared graph-spectral coordinates, whose signs define candidate regions and whose continuous values provide features for region-specific predictors. Tree dynamic programming selects region–predictor pairs to maximize workload coverage under memory and approximation-error budgets. On frozen Qwen3.5-9B, MLP Cache replaces 24–28% of MLP calls on MME-Coarse and MMStar-Coarse, with accuracy decreases of 1–2%. The cached model achieves a batch-one decoding wall-clock speedup on A100. Controlled MLP experiments show that cached prediction offers greater speedups for larger MLPs, the gains depends on the cost of uncached execution relative to cache lookup. For a 9B-shaped MLP, hit rates of 30% and 50% yield A100 prefill speedups of and , and M5 Pro prefill speedups of and .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.