Not All History Matters Equally: Tokenization through Instance-Adaptive Aggregation and Time-Decay Preservation
Abstract
Embedding quality constrains the achievable performance of time series forecasting. Patch-based tokenization preserves local temporal structure,but assigns a uniform token to every patch regardless of its predictive value, leaving the downstream model to filter irrelevant information and resolve redundant patterns. We propose RePatch, a forecasting tokenizer built on the principle that not all history matters equally. Through instance-adaptive aggregation, it learns input-dependent weights to suppress uninformative patches and combine related patterns across non-contiguous regions into a compact set of tokens. The forecasting objective jointly trains the tokenizer and downstream model, directing the limited token budget toward information useful for prediction. The learned aggregation also exhibits a preference for recent history.To reinforce this behavior, we introduce a time-decay prior that merges recent patches with the learned tokens. An input-conditioned allocatorcontrols their mixing weights, with stronger preservation for patches closer to the forecast and progressively weaker preservation for more distant candidates. This fusion retains recent context alongside global patterns without increasing the number of tokens. Search-based comparisons on five benchmarks show lower average MAE than the compared methods. Aggregation comparisons, recent-candidate ablations, and parameter sweeps further examine the design and its sensitivity to tokenization choices.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.