acceptodds
Under review as a conference paper at ICLR 2027

KlaraBench: Benchmarking Memory Between the Lines for AI Agents

Abstract

Long-term personalized agents need not only to remember facts and preferences explicitly expressed by users, but also to infer individual-level behavioral representations from past behavior that can transfer to new scenarios. Recent benchmarks evaluate cross-record integration and behavioral patterns, but few directly assess whether models can transfer user representations induced from behavioral evidence to novel scenarios under strict semantic isolation. To this end, we propose , a memory benchmark grounded in the Big Five personality framework that evaluates cross-scenario behavioral trait induction and transfer. tests whether memory systems can infer stable user characteristics from behavioral evidence across multiple unrelated scenarios and apply them to novel contexts. We evaluate four base language models and four external memory systems, with Profile Induction as a profile-based aggregation reference. Under the unified evaluation protocol, the tested retrieval-based memory systems do not consistently outperform full-history input, with performance varying substantially across base models. Counterfactual analysis shows that responses generated by external memory systems exhibit weak persona-conditioned specificity. Scenario-isolation interventions reveal that model performance improves substantially when the history contains the same decision object as the test request, indicating that strict scenario isolation removes a local matching shortcut and makes the task more dependent on cross-scenario behavioral abstraction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.