Distilling Across Tasks for Generalizable Memory Policy in Personalized Agents
Abstract
Personalized agents must use a user’s history across changing tasks and domains. Existing memory systems typically rely on hand-designed routines or task-specific training, leaving open how to learn a shared, transferable memory policy. We introduce MEMORYOPD (Memory On-Policy Distillation), which trains a small student across tasks using the token distributions of specialist teachers operating within a memory harness. On MemoryCD, this harness combines episodic evidence with a semantic user profile; distillation transfers the teachers’ memory-conditioned behavior on the student’s own generations, providing dense token-level supervision without a downstream task reward. Our analysis characterizes a geometric teacher consensus at a common context and gives sufficient conditions for a memory-conditioned policy to transfer across domains. We evaluate generalization through two complementary settings with separately trained policies. A policy trained jointly on MemoryCD tasks is reused without parameter updates on eight held-out product domains. Separately, a policy trained on LongMemEval is transferred to LoCoMo without further training. In domain, MEMORYOPD configurations outperform the seven comparison methods on all eight personalization metrics for both student families, including a 27% reduction in Qwen rating MAE (0.490 to 0.357). The transferred MemoryCD policies lead the eight-domain averages on every metric; the LongMemEval policy reaches 0.695 held-out accuracy (best baseline 0.629) and transfers successfully to LoCoMo. These results support learning memory behavior that balances multiple tasks and remains useful beyond its training distribution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.