acceptodds
Under review as a conference paper at ICLR 2027

Scaling Laws for Order Flow Generation

Abstract

Limit order books capture financial markets at their most granular level, and autore- gressive generative models have emerged as a powerful approach to modeling their dynamics, yet their scaling behavior remains poorly understood. We study how the size of state-space models (Mamba-3) and cumulative training-token exposure affect cross-entropy (CE) on S&P 500 order flow, examining both logged training CE from 2022–2025 and forward-time test CE on January 2026 data. We analyse 32 training lineages across 12 model sizes, spanning over two orders of magnitude (2.63–293.28 million parameters), with 249 January checkpoints evaluated on the same 487 tickers. Under the realised training budgets, January endpoint CE improves sharply across the smallest models and then varies non-monotonically with size. An additive parameter–exposure law predicts held-out lineages and sizes within the observed range more accurately when restrictive exponent bounds are relaxed. Most whole-size prediction error comes from extrapolating to the smallest model, while coefficient estimates remain sensitive to size coverage and fitting choices. We also evaluate generated continuations with LOB-Bench, comparing observed and generated event and book-state distributions after processing gener- ated events through a matching engine. We relate the resulting discrepancy scores to CE at matching checkpoints. Using activity-derived hardware-cost estimates, we compare IsoFLOP quadratics with the additive surface at fixed compute budgets. The results show where the additive law describes order-flow prediction and where its fitted exponents depend on the experiment and fitting choices.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.