Decentralized Online Learning with Bandit Feedback in Fair Value Games
Abstract
We study decentralized online learning in task-based Fair Value Games (FVGs), where heterogeneous agents repeatedly select tasks and receive their marginal contributions to unknown, identity-dependent coalition values. We introduce OSD-FW, a projection-free Frank–Wolfe method that uses one sample per update and adapts a known template for non-convex objectives. Its new component is a sampled drift correction. The centralized form of that correction needs the joint change in every agent's policy, which no individual agent holds, and earlier decentralized methods substitute a compact summary statistic that identity-dependent values do not admit. Ours is built instead from single actions that agents draw and announce for themselves, at a per-round query cost linear in the numbers of agents and tasks. Because the Frank–Wolfe gap equals total exploitability in these games, the resulting stationarity bounds are Nash-regret bounds, with under unilateral-oracle feedback, once the correction is added, and under pure bandit feedback, the last slower than the best rate known for this class. Under a matched value-oracle budget, the drift-corrected variant ranks fourth on average Nash gap but attains the smallest final-iterate gap among the decentralized methods tested.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.