acceptodds
Under review as a conference paper at ICLR 2027

Memory-Use Policy Collapse in Role-Playing LLMs: Mitigation with Diversity-Controlled Post-Training

Abstract

Memory use in role-playing dialogue is inherently one-to-many, as the same scenario allows multiple valid memory-use policies. However, models may concentrate on only a few of these policies, producing repetitive responses that may weaken user immersion. We identify this phenomenon as Memory-Use Policy Collapse, which occurs in memory selection and memory use. We find this problem in strong API models, showing that valid memory use does not guarantee diverse memory use. To mitigate it, we introduce a plug-in memory planner between retrieval and response generation. The planner produces a memory-use plan that specifies which memories to use and how they should shape the next response. We use instance-specific rubrics to construct multiple valid plans with distinct plan-intent types for each scenario, and then train the planner with Diversified Supervised Fine-Tuning (D-SFT) followed by Diversity-Controlled GRPO (DC-GRPO). This training aims to improve plan validity while maintaining diversity. On the test set, DC-GRPO achieves a plan validity rate of 75.34%, compared with 60.16% for the vanilla SFT baseline. DC-GRPO increases Effective Plan-Intent Types@10 from 1.649 with the vanilla GRPO baseline to 4.978, achieving greater diversity among valid plans. Our planner also improves the validity and diversity of generated responses. Human evaluation further shows an overall preference for responses generated with DC-GRPO.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.