BEST-OF-BOTH-WORLDS CAUSAL BANDITS WITH MEDIATOR FEEDBACK
Abstract
We study a bandit setting described by a causal graph , where is an action, is an observable mediating state separating from the loss , and is an unobserved variable. The setting has been proposed and studied independently in the causality literature under the name of *causal bandits with post-action context*, and in the classical bandit literature under the name of *bandits with mediator feedback*. Similar to both bodies of literature, we assume that the mapping is fixed. We propose a best-of-both-worlds algorithm, *Causal Tsallis-INF*, which provides a worst-case regret guarantee in case of adversarial and instance-dependent regret guarantees for stochastic , as well as other forms of for which the regret satisfies the self-bounding constraint of Zimmert, Seldin (2021). Our setting matches the one studied by Eldowa et al. (2024) in the classical bandit literature and generalizes the setting studied in the causal bandit literature, where is assumed to be stochastic. We advance both strands of literature. Our stochastic regret guarantee improves the logarithmic dependence on the horizon in the EXP4 analysis of Eldowa et al. (2024) and adapts to both action and attainable mediator complexity. Our adversarial guarantee also extends beyond the stochastic assumptions underlying the causal bandit methods. We further provide an empirical evaluation demonstrating that for stochastic our algorithm performs comparably to state-of-the-art Causal Thompson Sampling and Causal UCB2 algorithms, and that, in contrast to our algorithm, the latter are not robust to non-stochastic , at the same time, our algorithm consistently outperforms the EXP4 algorithm in all the experiments. References: Eldowa et al. (2024): Information capacity regret bounds for bandits with mediator feedback. JMLR Zimmert, Seldin (2021): Tsallis-INF: An Optimal Algorithm for Stochastic and Adversarial Bandits , JMLR
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.