SIRA: Self-Improving Recommendation Agents under Demand Shifts
Abstract
Conversational agents increasingly sell products and services, and to be profitable, an agent must (i) find out what each customer wants, (ii) offer what the customer will buy, and (iii) complete the purchase before the customer loses patience. No benchmark measures profit in sales conversations that require all three. We introduce \env, a benchmark that scores the profit an airline keeps when an agent sells flight add-ons to simulated customers. Each customer has a hidden type that determines what they buy, so the agent must learn which questions reveal the type and what each type buys as demand shifts. Each turn, the agent makes one of three tool calls, to ask a question, offer an add-on, or complete the purchase. Unlike language-model customers, which can be persuaded, \env customers react only to these calls, making results reproducible and the best profit exactly computable. Neither prompting a language model nor fine-tuning it with standard recipes makes it a profitable seller. A hand-written sales strategy does better by asking for the facts each add-on requires and keeping a turn to complete the purchase, but ranks offers by revenue rather than profit. We propose \method (Self-Improving Recommendation Agent), which imitates this strategy, then corrects its objective with reinforcement learning on the profit of its own conversations. \method (i) earns 110.7% more profit than the prompted model, (ii) beats the strategy it imitates in every demand phase and by 14.5% overall, with the same test-time information, and (iii) learns which add-ons earn more, coming within 3.4% of the same strategy given product costs. Yet an agent that knows how customers are generated but not their price limits could capture at least 68% of the add-on profit of an agent with this privileged information, and \method captures 37%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.