Learning What to Practice Next: Experience-Adaptive Self-Evolution for Multi-Turn Empathetic Dialogue
Abstract
Reinforcement learning with verifiable feedback typically updates model parameters while leaving the distribution of subsequent experience fixed. This separation is especially limiting in multi-turn empathetic dialogue, where interaction conditions shape entire trajectories and their training value changes as the policy improves. We introduce experience-adaptive self-evolution, which turns verified emotion outcomes into an online signal for deciding what the policy should practice next. Continuous outcomes train the dialogue policy, while thresholded group outcomes estimate the training value of reusable intent-state units and reallocate future interactions. Hierarchical evidence sharing stabilizes sparse estimates, and uncertainty-guided exploration with uniform rehearsal preserves coverage. With the simulator and verifier frozen, the framework adapts experience without changing the role-playing target or increasing the rollout budget. On SAGE, it raises Qwen3-8B Overall from 53.87 to 79.24, outperforming protocol-matched uniform emotion-reward RL by 7.23 points; the advantage also holds at Qwen3-4B by 4.14 points. Cross-benchmark and human evaluations further show improvement beyond the training-aligned metric.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.