Learned Prefetching with Delayed Utility
Abstract
While learned sequence models have proven highly effective at capturing complex memory access patterns, they are conventionally framed as short-horizon predictors and evaluated under idealized, instant-fill assumptions. The rapid shift toward disaggregated memory topologies—where memory is pooled across high-latency fabrics—fundamentally breaks this paradigm. In these environments, higher and variable latencies mean that timeliness dictates system performance: a highly accurate prediction that arrives too late fails to hide latency, while an early prediction wastes limited bandwidth and pollutes the cache. To address these disaggregated constraints, we argue that prefetching must be recast from next-step prediction to a delayed-utility learning problem. We propose a multi-horizon sequence model that predicts across several future lead times, exposing a richer set of candidate accesses tailored to the deployment's specific latency regime. Evaluated under a rigorous, stall-aware framework using real-world memory access traces, we demonstrate that standard single-horizon models frequently fail in high-latency environments, whereas our multi-horizon approach successfully captures the deployment-dependent utility of different lead times at a modest inference cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.