When Delayed Feedback Changes the Objective: Benchmark Validity in Sequential Decision Making
Abstract
Modern recommender, advertising, and online experimentation systems often act several times before a consequential outcome matures. Delayed-learning theory usually treats this lag as an information constraint while taking the evaluation loss as given. We study the prior question: when does the route from delayed observations to action-level losses define a valid benchmark? We separate structural decision time from observation time and show that validity is governed by preservation of action gaps, not by delay length. For any realized action sequence, route-level performance transfers to the structural benchmark up to the cumulative distortion of action gaps. When structural margins are present, this gap-control certificate also recovers rankings and optimal actions; when persistent conflict crosses those margins, a learner may optimize the routed objective while accumulating structural regret. We further prove that fully unlabelled delayed observations can leave route validity non-identified even under adaptive data collection, motivating evidence-qualified audits rather than assumption-free corrections. Controlled experiments stress-test the predicted transfer and failure regimes; logged-data studies diagnose attribution sensitivity and illustrate the distinction between predictive fit and decision recovery on observed support. Reliable delayed-feedback learning therefore requires two certificates: optimization of the operational route and validation of the benchmark it induces.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.