RosKV: Robust Spectral Key-Value Cache Compression for Time-Series Foundation Models
Abstract
Autoregressive time-series foundation models (TSFMs) enable zero-shot long-horizon forecasting, but real-world deployment on resource-constrained hardware is bottlenecked by a growing key–value (KV) cache. While caching past states eliminates quadratic recomputation, persistent memory scales linearly with look-back length and horizon. Furthermore, TSFM prefill importance is delocalized in time yet localized in frequency: seasonal cycles span many look-back tokens, so token eviction and uniform time-domain quantization are ill-matched, whereas a small fraction of RFFT bins carries most K/V energy. We present , a memory-first, training-free spectral compressor for frozen TSFMs designed for memory-constrained deployments where persistent cache capacity is the primary barrier. RosKV (i) applies a fixed head-dimension rotation so NF4/NF2 block scales are not dominated by outliers on the few kept bins, (ii) transforms along the look-back with an RFFT, (iii) filters K and V separately by spectral energy, and (iv) stores adaptive mixed-precision coefficients, reconstructing an approximate dense prefill at every decode step before attention. Across nine datasets, four horizons, and four backbones (TimeMoE, Chronos, Moirai, TimesFM), RosKV consistently outperforms NLP token eviction (SnapKV, PyramidKV), coordinate quantization (KIVI), and frequency baselines (FreqKV), staying within of Full-Cache aggregate MSE (with matching rankings confirmed on GIFT-Eval). On dense-token architectures where persistent KV storage forms a critical inference bottleneck (TimeMoE), RosKV slashes persistent cache memory by on average ( MB vs. MB) and up to at short horizons ( MB vs. MB), while maintaining competitive decoding latency ( ms vs. No-Cache ms).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.