acceptodds
Under review as a conference paper at ICLR 2027

When Biased Offline Data Help Online Learning: The Role of Action Contrasts

Abstract

Historical reward levels can be systematically biased while preserving the action comparisons needed for a new decision problem. However, guarantees based on absolute reward-bias bounds can overlook this benefit when all actions share the same reward-level shift. We isolate this information object in a finite-cell Bernoulli contextual bandit by bounding the diameter of source–target reward biases across actions and leaving their common level unrestricted. With a supplied valid relative-bias bound , context cells, actions, balanced source observations, and target rounds, upper and same-model minimax lower bounds share the rate up to logarithms. For fixed , the minimax regret is if and only if and . Unlike absolute-discrepancy models, this class permits a large action-common shift; the familiar minimum-of-two rate is characterized here for that quotient information class. A relative-confidence allocation turns historical contrasts into observable gap bounds and permits any predictable proposal; information-directed sampling is one instantiation, not a condition for validity. When must be learned, target calibration is charged to regret, and a honesty barrier prevents this route from uniformly recovering the strict order gain within the same horizon. Controlled experiments recover the predicted phase boundary, while response-bank experiments expose the finite-sample costs of protection and of learning the contrast certificate.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.