The Window Lever: Observation-Window Length Outweighs Observed Scale and Architecture Effects in Customer Behavior Forecasting
Abstract
Forecasting a customer's next year: add customers, grow the model, switch architectures, or look further back? We test all four, one knob at a time. Looking back wins. Stretching the observation window from 8 to 16 weeks adds to log-space in all eight models at 1M, and the simplest gains most. Test-set composition is the obvious suspect, so we fix test customers and targets across windows (training cohorts still grow with the window). The gain gets bigger, not smaller, and the matched transformer gain beats all 1,000 permutation-null draws. The direction replicates on a public dataset. For sequence models, the other levers move an order of magnitude less. the customers shifts by at most , as does a larger transformer (three-seed mean), and on one seed, five sequence architectures land within at a matched objective at 16 weeks. All three are one-knob contrasts in small-model recipes; joint scaling is untested. We internally registered four predictions, with falsifiers, before the deciding evaluations. Three held. The one fully blind prediction failed: an activity-mixture reading called for from 24 to 32 weeks; we measured . The pooled weekly-aggregate probe plateaus near 22 to 24 weeks, but activity-specific refits keep climbing: no information ceiling is established. One limit remains: every test anchor falls in one eight-week band, so window length and calendar coverage move together. The direction is robust; the magnitude is not identified. Observation-window length belongs beside model selection as a first-class design axis: what a sequence model is shown matters more here than what it is.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.