acceptodds
Under review as a conference paper at ICLR 2027

Shared-Observation Chaining for Hierarchical Adversarial Noisy Bandits

Abstract

We study finite-time regret in adversarial noisy bandits with bounded conditional-mean feedback. Fixed-scale optimizer-cover scalars are not rate-complete: two finite-action classes agree on a nontrivial scale range yet have exponentially different regret. On a product routing model with templates, a known-horizon -time learner uses one route/reward observation to update all active levels and attains expected regret against adaptive nonanticipating environments; an oblivious Bernoulli subclass matches this rate. Even the optimistic fixed-rate one-global-Shannon ExO expression, with arbitrary finite real-valued Borel scores, is . A simultaneous product packing matches the complete active-coordinate finite-horizon profile ; in the information-limited regime this becomes , without a block-allocation depth loss. Under explicit margin, finite-range, matching, quadratic-regime, allocation, saturation, and terminal-control conditions, general optimizer hierarchies admit a conditional chain–packing comparison.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.