BIND BEFORE YOU BUYCANDIDATE-CONSISTENT CONSTRAINT LEARNINGFOR INTERACTIVE SHOPPING AGENTS
Abstract
Interactive shopping agents must find a single product configuration that satisfies all user requirements. Finding a black shoe and finding a size-42 shoe, however, does not mean finding a black shoe in size 42. Evidence from different products or incompatible options can appear collectively sufficient even when no single configuration satisfies the request. Meanwhile, terminal rewards provide little guidance on which intermediate decisions made valid progress. We introduce ShopBind, a candidate-consistent constraint learning framework that connects supervised fine-tuning (SFT) and reinforcement learning (RL) through a shared requirement-to-configuration binding state. For each observed configuration, this state records whether each visible requirement is supported, contradicted, or unresolved. During SFT, the agent learns to generate a compact binding trace before its next action, making product-specific constraint reasoning part of the policy. During RL, the binding state is independently reconstructed from visible interaction prefixes to identify first-time expansions of contradiction-free requirement coverage. This progress signal distinguishes trajectories with tied terminal outcomes and redistributes a bounded share of trajectory advantage toward the steps responsible for valid progress, while preserving the priority of distinct terminal outcomes. ShopBind leaves the benchmark reward, action space, rollout budget, and number of policy calls unchanged. On 1,459 Single-Turn tasks in ShopSimulator, ShopBind with Qwen3-8B achieves 45.92% full success, with loose and strict constraint scores of 65.76 and 50.00, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.