acceptodds
Under review as a conference paper at ICLR 2027

RSCB: REGIONAL SEMANTIC CONTEXTUAL BANDIT FOR OPTIMAL SELECTION OVER LLM-GENERATED DYNAMIC CANDIDATES

Abstract

Large Language Models (LLMs) increasingly generate multiple candidate solutions for complex tasks such as bug localization, code repair, and tool selection. Selecting among these candidates remains challenging because the available arm set changes across rounds and newly generated arms may have no interaction history. Identity-based or disjoint contextual bandits cannot transfer feedback to such arms, while globally shared semantic linear bandits may suffer negative transfer when heterogeneous context-arm relationships cannot be captured by a single reward parameter. We propose Regional Semantic Contextual Bandits (RSCB), which maintains independent linear reward models over semantic interaction regions and routes each context-arm pair to its corresponding regional model. RSCB selects arms using region-specific upper-confidence scores, transferring feedback among semantically related pairs within each region while restricting transfer across heterogeneous regions. For semantic regions, we prove a regret bound of under regional linear realizability and correct fixed routing. Consequently, for fixed , RSCB retains the regret rate independent of the number of distinct arms. We further establish an lower bound for arm-independent learning. Finally, under a semantic-continuity assumption, we bound the prediction uncertainty of unseen arms within a correctly routed region. We evaluate RSCB on SWE-bench and an Android/Java security bug dataset, using LLM-generated root-cause hypotheses as arms and objective rewards derived from ground-truth patches. On held-out bugs, RSCB reduces cumulative regret compared to SC-LinUCB and kernelized baselines, with statistically significant improvements on the security dataset. Additionally, RSCB also learns a multi-region structure () on the public BugsInPy benchmark, achieving the lowest regret among all methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.