PolyEvent: A Large-Scale Benchmark for Inductive Relational Event Forecasting
Abstract
How should we evaluate forecasters on unseen event groups whose markets have trading histories? Reviewed benchmarks do not report this comparison. We introduce PolyEvent, an open event-sequence benchmark using Polymarket's official event groups. It aggregates 1.03B order-book fills into 281M (market, block) events across 652,016 groups; experiments use a fixed 10M-event subset. Tasks are binned block-gap prediction and three-way price-change prediction (UP, FLAT, DOWN); the relational track evaluates the latter. Separately trained own-history and context-augmented models are frozen and compared on identical targets from held-out multi-market groups, using strictly earlier-block information. Scores average over events (event-micro) or markets (market-macro). Construction choices—fill-level events, timestamp order, order-direction marks and calendar splitting—each change what is evaluated. As a use case, context lowers negative log-likelihood on a sealed whole-group holdout (primary market-macro ΔNLL = −0.0490 nats), exceeding the 0.010-nat margin fixed before any sealed prediction. On development data, an audited placebo using other groups with matched sibling count yields ΔNLL = −0.0018 nats, its interval containing zero; it does not isolate event-group identity. In a post-hoc study of 156 development threshold groups, features encoding threshold order improve over pooling and a capacity-matched shuffled-role control (market-macro ΔNLL = −0.0250 nats against the control, simultaneous 95% CI [−0.0377,−0.0122]); event-micro gains are smaller. Data and code links will follow publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.