acceptodds
Under review as a conference paper at ICLR 2027

Sharp Safety and Critical Expectation Defects in Softmax Gradient Bandits

Abstract

Softmax gradient bandits adapt their exploration through updates driven by sampled rewards. Choosing the learning rate therefore requires understanding how these updates shape cumulative regret. Almost-sure convergence alone leaves this question unresolved, because rare trajectories can retain substantial expected loss even as typical trajectories learn the optimal arm. For IID arms with a unique optimum, a common predictable baseline, reward-residual bound , and gap floor , we identify the exact learning-rate ceiling for class-uniform logarithmic expected regret, including the endpoint. Below this ceiling, every fixed positive rate yields cumulative pulls to leading order for each suboptimal arm with gap , almost surely and in . Thus typical and expected allocations agree to leading order, with equal leading regret contributions from all suboptimal arms. At the ceiling, an explicit two-arm family with signed rewards retains logarithmic expected regret, but rare, prolonged excursions toward the suboptimal arm raise its leading coefficient above the typical-path value. A positive harmonic function quantifies this excess and its dependence on initialization. A uniform near-critical expansion for the same family connects the critical correction to polynomial expected regret above the ceiling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.