Finding Epsilon-Best Arms in Structured Bandits via Fused Feedback
Abstract
To model real-world scenarios where a learner receives diverse, potentially biased feedback signals, we introduce a structured multi-armed bandit framework with unaligned fused feedback. In this setting, at each round, the learner can either select a single arm to observe its standard reward feedback or select a pair of arms to observe the corresponding dueling feedback. Crucially, each feedback modality has its own latent parameter: the expected feedback associated with an arm is the inner product of its feature vector and the corresponding modality-specific parameter. We propose Wit-BotElim, an algorithm tailored for -best arm identification under fused feedback, which leverages a novel fused-elimination rule to efficiently aggregate disparate feedback modalities. By proving a generic sample complexity lower bound, we demonstrate that Wit-BotElim achieves near-optimal performance in key practical regimes. Finally, numerical experiments indicate the superiority of Wit-BotElim over shared-parameter methods under latent-parameter mismatch, which is attributed to the delicately-designed fused-elimination.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.