acceptodds
Under review as a conference paper at ICLR 2027

Hora: Learning a Universal 128-Word Vocabulary from Billions of Cross-Domain Time Points for Zero-Shot Forecasting

Abstract

Time series foundation models (TSFMs) are pretrained on large cross-domain corpora to forecast unseen series in a zero-shot manner. They inherit many designs from large language models (LLMs). An LLM is built on a vocabulary, but time series have no native vocabulary. This raises a fundamental question: *what should the vocabulary of time series be?* Existing approaches either have no vocabulary, borrow one from language, or learn one that collapses similar sequences. To answer this question, we propose **Hora**, a TSFM built on a **temporal vocabulary** of 128 words learned from billions of cross-domain time points. Hora *reads* each patch as a mixture of the words, whose weights constitute its *temporal grammar*. It *consults* the vocabulary in every layer as a global memory, and *writes* the forecast from the same words. Experiments on GIFT-Eval show that Hora achieves state-of-the-art zero-shot forecasting performance. Visualizations further demonstrate that the interpretable temporal grammar expresses patterns shared across domains. Our code is available at https://anonymous.4open.science/r/Hora.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.