A General Framework for Budgeted Threshold Incentives on Request
Abstract
On-demand delivery platforms pay riders through incentive activities whose tiers are set from recent completions of riders with a similar history, so that a little extra effort earns a clearly stated reward. Operators request such plans for changing periods, rider populations, payment rules and budgets, days ahead and within minutes, often for holidays or bad weather where randomized trials are scarce and take months to collect. We present a request-driven framework that composes four stages—conditional prediction, population reduction, trajectory integration and budget allocation—through seven replaceable modules that exchange conditional trajectory laws, whose award probabilities and award-marked moments give payment and uplift for any activity rule. To shorten a long randomized campaign, a response-correction step reweights trajectories from abundant no-offer history to match the moments of a short pilot. We prove that, on a fixed plan menu and given the stage errors, the end-to-end value loss is bounded by the sum of four stage terms—synthetic data, moment matching, integration and decision—with constants that cannot be improved from the final tables, and that for every stage there are instances on which omitting it leaves an error floor the others cannot remove. On 3,000 riders over 45 weekly origins, re-drawing trajectories and solving exactly for every request takes 117.8 s per week-long request against 3.07 s for the framework; including the one-off sampling pass, all 127 windows of a week are answered 11.04× faster with identical scenarios and at most 0.92% value lost by the allocation on the audited week-long tables, a gain that comes entirely from reuse, while a point forecast, independent days or an equal budget split each lose accuracy. Holding cities out of training changes plan-cost error by +0.003 [−0.018, +0.020], within a pre-set margin. On 24 new controlled response laws, with the plan, the optimum and validity all judged within 20% of the budget, the response correction with a one-week pilot lowers regret by 51.2% relative to a trial with the same nominal randomized rider-weeks and by 53.6% in the stricter pre-specified 10% band, and in the exact-summation analysis of these week-long requests a four-week pilot, with 1/4.5 of the randomized rider-weeks, comes within +0.007 [−0.012, +0.027] of an 18-week trial. In the registered primary conditions of two studies where windows, populations, rules and binding budgets change from request to request, the framework's regret is below that of a trial with the same nominal rider-weeks and below that of dose interpolation of the same pilot data at every pilot length, and reusing its one-off preparation answers 60 requests 14.1× and 2.70× faster than re-running the pipeline for each, with identical answers. When the paid window or a selected subpopulation changes behaviour, a window-aware correction lowers one-week regret by 0.195 against direct reuse. Against a nine-offer trial fitted with the framework's own dose curve, one-week regret is 0.055 lower; 73.5% of the gain over per-offer fits of that trial comes from sharing the tilt across offers. Public retail data and three payment rules confirm the transfer with identical exact outputs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.