acceptodds
Under review as a conference paper at ICLR 2027

Fast A/B-Testing: Efficient Policy Evaluation via Tree-Coupled Feedback Sharing

Abstract

Online platforms increasingly compare many adaptive decision policies—ranking systems, recommendation algorithms, pricing rules, and language-model agents—while each reward-bearing interaction can be costly or risky. The standard recipe is a direct A/B/n test, which runs each of candidate policies over an independent -horizon and therefore uses outcomes. We introduce Tree-Coupled A/B Testing (TCAB), an exact feedback-sharing design for reducing the outcome usage while preserving the marginal trajectories. At a high level, the policies are associated to a tree structure and are executed by any topological order. The observed outcome is shared along an edge when the two policies on this edge can be coupled. Every policy retains exactly its standalone finite-horizon trajectory law, even though the policies are deliberately dependent. If records a mismatch on tree edge at round , the number of outcomes satisfies the pathwise identity . This cost is conditionally optimal for the selected tree, and we propose a heuristic for empirically finding the optimal tree. For fixed , sublinear pseudo-regret of every policy can imply . We also obtain finite-sample variance bounds for pairwise policy contrasts. Experiments on language-model evaluation and adaptive bandit policies demonstrate substantial improvements—around 20-50% outcome save and improved policy evaluation precision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.