Direct Contextual UCB under Exact Cauchy Noise
Abstract
Contextual bandits whose reward noise is so heavy-tailed that the mean does not exist break the usual pseudo-regret benchmark; we study exact symmetric Cauchy noise, targeting latent location regret. We analyze Direct-Cauchy-UCB, a one-pass arm-local projected gradient update driven by the Cauchy likelihood score, with optimism from the current iterate, and state per arm. Under bounded parameters, bounded contexts and directional excitation — every direction excited at every step—it attains high-probability location regret . Our lower bound of order matches this up to logarithmic factors once the horizon is large enough. The same update keeps that rate under a weaker windowed persistent-excitation assumption, covering deterministic adversarial context designs. Without any excitation assumption, an ellipsoidal variant is sublinear at with per-arm state. Beyond exact Cauchy noise, a sign (median-score) variant handles symmetric heavy tails at the same rate, with no moment assumption, given a known lower bound on the noise law's slope at its median. Bounded per-arm deviations of size from Cauchy cost the exact-score policy an additive , against a floor proved for that class. On synthetic exact-Cauchy environments our policy attains the lowest location regret among a strengthened suite of robust baselines at low and moderate noise scales.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.