Moreva: When Does Attention Help Mamba Forecast Event Streams?
Abstract
Event streams often comprise irregularly timed events with categorical types, motivating joint prediction of the next event's time and type. State-space models such as Mamba efficiently summarize event histories in a fixed-size recurrent state, but may lose information about specific distant events. This limitation motivates hybrids that augment recurrent updates with attention-based retrieval. Although such hybrids have shown promise across sequence-modeling tasks such as Large Language Model and Time Series Forecasting, when attention improves temporal event forecasting and whether standard benchmarks contain dependencies that require retrieval beyond a recurrent summary remains unclear. We investigate these questions with Moreva, a shared event-history model that adds parallel causal-attention branches to selected layers of a pretrained Mamba backbone. On a synthetic process with known dependency length, attention lets Moreva recover type information placed hundreds to thousands of events back, which the Mamba-only model and the baselines miss. On EasyTPP benchmark datasets, Moreva achieves the state-of-the-art compared to time-point-process, regular time series and language-model-based baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.