acceptodds
Under review as a conference paper at ICLR 2027

Efficient Zero-shot Spatio-Temporal Forecasting via Multi-scale Expert Distilled Mixed Hashing Attention

Abstract

Zero-shot spatio-temporal forecasting is crucial for applications such as smart cities and IoT, and while Large Language Models (LLMs) show promising capability in this setting, their high computational cost limits real-time deployment. Designing an efficient student model to distill knowledge from LLMs offers a promising solution, though it remains challenging. In this paper, we propose Multi-scale Expert Distilled Mixed Hashing Attention (MSED-MHA). We first employ Linear Attention to capture global spatio-temporal trends with high efficiency. In addition, we introduce Spectral Hashing Attention to model local dependencies through hashing-based Top-k retrieval over relevant neighbors. To address the intractable optimization with multiple hash constraints, we adopt a spectral learning approach to enable accurate and efficient retrieval. In addition, a selective Multi-Scale Expert Distillation guides the student across temporal scales, applying teacher supervision only when beneficial, thereby mitigating error propagation. Experiments on traffic, finance, and weather benchmarks demonstrate that MSED-MHA achieves competitive forecasting performance with high computational efficiency compared to state-of-the-art methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.