ECB: Evidence-Conditioned Continuous Bandit for Recommendation Agents
Abstract
Recommendation agents must decide where to retrieve evidence, how much evidence to acquire, and when further exploration is no longer worthwhile. This creates two fundamental challenges. First, reranking an initial candidate list cannot recover items that were never retrieved, making candidate coverage dependent on the retrieval workflow itself. Second, retrieval and candidate scoring consume limited resources, so exploration must balance the discovery of new items against its additional cost. We propose Evidence-Conditioned Continuous Bandit (ECB), a budgeted controller for adaptive workflow control. ECB introduces a gate-before-scoring mechanism that first examines the base recommender's output and determines whether further evidence should be acquired. When exploration is triggered, an Evidence-aware Ranker integrates user history, contextual information, item attributes, and construction evidence to produce candidate-level retrieval signals. These signals condition a joint continuous policy that controls interest focus, evidence mixing, and retrieval-source allocation. The policy therefore coordinates what to retrieve, how broadly to explore, and when to stop, while the base recommender ranks the expanded candidate pool. ECB learns this policy offline from sampled workflow trajectories using a target-conditioned proxy reward that jointly accounts for recommendation outcomes and resource costs. A likelihood-ratio update enables policy learning through the discrete retrieval executor. Experiments on multiple recommendation benchmarks show that ECB improves recommendation quality and target coverage, demonstrating the value of evidence-conditioned exploration and budget-aware workflow control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.