Learn to Sell, Not What to Sell: Dualistic Self-Evolution (DSE) of LLM Sales Agents with Human Nature as Reward (HNR)
Abstract
Text-based selling is a long-horizon decision process in which customer behavior evolves across interactions and damaged trust is rarely restored. Delayed sales outcomes complicate autonomous improvement, so humans must still supply dense rewards and evaluation. Motivated by the distinction between transferable sales skills and business-specific knowledge, we introduce Dualistic Self-Evolution (DSE) with Human Nature as Reward (HNR), which separates a profession learned in the model weights from business knowledge, rules and tools carried by an external harness. The inner loop learns from open cross-category dialogues and 8 business-agnostic prototype tools through supervised fine-tuning followed by reinforcement learning. HNR is drawn from sales practice: it rewards selling with the grain of the customer (answer first, verify before claiming, move one step forward), penalizes the model's own counterproductive habits, and assigns zero episode reward to unsupported claims and other credibility violations. Outcome-gated turn-level scoring reduces the share of rollout groups with zero reward variance from roughly two thirds under outcome-only rewards to 3–5%. The outer loop updates the harness through reward-guided edits validated by real-conversation replay. The loops share no parameters and have separate reward agents and human arbiters: senior salespeople assess craft and quality-control staff assess factual and procedural correctness. With the target business excluded from training, frozen models transfer to a production used-car harness that maps the 8 prototypes onto 16 business tools. On replay of 600 real customer turns with full conversation histories, the trained 9B model improves strict response quality by 8.3 percentage points over its same-size base (13% relative; three seeds; paired 95% CI [+4.4, +12.0]), and the share of turns with a tool call rises from 57% to 80%; the 8B gain is smaller and statistically unresolved. On the out-of-domain Ψ-Bench, persuasion improves by 88% at 8B and 54% at 9B. At 9B the prototype-tool SFT carries the strict gain and reinforcement learning adds to it only at 8B; the gate's measurable contribution is gradient availability rather than replay score. Fine-tuning on business transcripts degrades replay quality and removes tool use, whereas the harness alone lifts untrained models above the deployed system and accommodates business changes without retraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.