acceptodds
Under review as a conference paper at ICLR 2027

The Price of a Fixed Scale: Exact Regret Constants and Adaptive Scaling for Tsallis-INF

Abstract

Best of both worlds bandit algorithms are usually compared by regret rates, but their stochastic leading constants determine whether robustness preserves asymptotic efficiency. We resolve this question for reduced variance -Tsallis-INF. For every fixed scale and fixed , its Bernoulli small gap coefficient is exactly under an explicit admissibility condition. The optimum is , strictly above the local Lai–Robbins constant , while the standard choice yields . Thus every fixed scale under this limit misses local Lai–Robbins efficiency. The formula points to local variance as the missing quantity. Guided by this decomposition, we adapt the scale using sparse probes that require neither instance parameters nor arm order. For two stochastic arms, a fixed clip family approaches for every interior mean, and a moving-clip learner attains on the central family under ordered limits. The proof identifies a harmonic occupation law and controls the singular regret cost using recovery estimates and inverse moments. The adaptive learners do not retain adversarial guarantees, and we do not provide a joint finite horizon and small gap rate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.