GCE: How much of a user's experience do language models actually use
Abstract
Persistent LLM assistants build up a long record of one user's decisions. Yet when such an assistant mispredicts a new decision, existing memory benchmarks cannot tell whether the history lacked the answer or the system failed to use it: they score against noisy labels, with no known bes t achievable loss. We introduce Generative Cognitive Evaluation (GCE), which measures how much of the predictive information in a user's experie nce a system actually uses. Each synthetic user is an executable preference model drawn from a declared finite family. Its choices are rendered as prose sessions that never reveal the schema, and exact enumeration gives the Bayesian posterior predictive for every held-out decision. This reference splits a system's loss into choice noise, uncertainty the history cannot resolve, and the system's own deviation, and defines an infor mation-use ratio : for the constant forecast , for exact Bayesian inference, and negative for worse than the constant. Across six preference families, the best of three frontier LLMs reading the full history (at low reasoning effort) realizes about three fifths of the available gain (), while the other two do no better than the constant ( and ). Summary, profile and retrieval memo ries do not close the gap, and two off-the-shelf memory systems add nothing beyond them: Mem0 behaves like episode retrieval, and Letta like in- context reading, since its agents rarely write memory. Using the LLM only to parse prose into features and fitting a simple statistical model re aches and on the two GPT backbones, close to logistic regression on the true features (). This suggests that the inform ation is recoverable from the prose and that much of the gap lies in how models aggregate evidence. GCE is a controlled diagnostic conditional o n the declared family, not a measure of performance on human preferences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.