Simulating Individual Delivery Riders from Their Own Histories with Measurement-Grounded Memory
Abstract
Delivery platforms want to simulate how an individual rider will behave by conditioning a language model on that rider's own history. When such a simulator must show LLM-written memory to its forecaster, the memory's numbers have to stay tied to the rider's measurements. We build the simulator around a measurement contract: source statistics keep their numerator, denominator, unit and window; LLM-written memory is stored as typed, unverified claims with declared dependencies; a deterministic check flags a claim only against a measurement of the same definition; and dependency closure retracts every statement built on a flagged claim. We prove that the retained memory is the unique largest conflict-free, dependency-closed sub-memory, that budgeted selection is feasible and maximal, and that a premise-preserving selector keeps both guarantees after retraction. Across 316 riders, conditioning on the rider's own history lowers offer-response log loss by 0.0892 relative to population context and raises accuracy by 15.30 and 8.98 points over wrong-identity and matched-neighbor histories, so the simulator represents the intended rider. An outcome-free audit of 590 memories for 295 riders finds 41.6% of the explicit numeric claims outside tolerance even when the memory is written from the very table the measurements come from; an independent reader of the same memories reproduces every parsed claim, source measurement and retraction set. In a preregistered two-week forecasting study with a Qwen3.5-9B writer and forecaster, typed memory with the consistency check is non-inferior and superior to the same memory shown as plain text on the 22 evaluable riders of a holdout group outside every earlier cohort, with differences of in offer log loss and in refusal-fraction error, and across all 295 riders it lowers the share of conflicting metrics on which the forecast lies closer to the stale claim than to the fresh measurement from 22.3% to 4.7%. A registered validation at five new forecast cutoffs finds checked memory non-inferior to plain memory on offer log loss, and on 789 never-exposed riders of a third sample checked memory lowers offer log loss relative to plain memory by 0.1239 with interval in a registered descriptive contrast. In a registered large-sample confirmation on a 9,350-rider cohort, combining an LLM forecast that reads the rider's memory documents and measurements with a smoothed personal statistical predictor lowers refusal-fraction error by 9% relative to that predictor alone on 1,998 confirmatory riders, , under both registered tests. The memory layer thus provides measurement-grounded maintenance with structural guarantees for simulators that must use external memory.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.