acceptodds
Under review as a conference paper at ICLR 2027

Policy Learning from Reward Totals

Abstract

A reward total records how much a group received, but not which interaction produced it. We study policy learning from interactions released as at most disjoint, unweighted reward totals. An encoder may use all contexts, actions, and logging propensities to form groups before seeing rewards; the learner receives the totals, metadata, and assignments. In a fixed two-policy family, data-independent grouping has worst-case regret at least of order , whereas two comparison-aware totals attain . For larger finite libraries, randomized block encoders trade the number of totals against estimation variance, and lower bounds cover every outcome-blind partition fitted to the complete metadata batch. The bounds match in specified polynomial-budget and large-sample linear-budget regimes. For policies generated by fixed gates and actions, a shared value-moment table requires at most totals per reporting window and supports optimization over all stochastic heads after release. The reward mean need not be realizable by the gates; under i.i.d. logging and coverage, the binary-action rate matches individual feedback. Simulations and a retrospective click-log study distinguish policy-selection error, same-log acquisition loss, and later-period return: shared totals reduce acquisition loss on the observed log, while later-period intervals do not establish return dominance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.