acceptodds
Under review as a conference paper at ICLR 2027

Retrieval Is Not Understanding: The Evidence-to-State Bottleneck in Personalized LLMs

Abstract

Long-term personalized large language models (LLMs) must use information from prior interactions to support decisions in the current context. Memory-centric systems have made substantial progress in storing, retrieving, and organizing historical information, while recent work has also introduced explicit user representations. However, it remains unclear how much user-state error persists when retrieval failure is explicitly removed. We isolate this post-retrieval problem as the Evidence-to-State bottleneck. Under an Oracle Evidence setting, task-relevant historical evidence is directly provided, allowing us to evaluate whether the model can recover the corresponding benchmark-defined user state without confounding retrieval errors. Using PersonaMem-v2, we represent each state by its type, subject, content, and applicability. Even under this favorable setting, a base LLM achieves only 19.81% Full-State Accuracy, indicating substantial residual error in user-state interpretation. To test whether this bottleneck can be substantially reduced with direct supervision, we apply a simple controlled state-supervision intervention, which raises Full-State Accuracy to 65.28%. We then perform matched state replacement while holding the query, Oracle Evidence, candidate options, and downstream decision model fixed. Replacing base-generated states with improved states raises personalized decision accuracy from 47.10% to 68.81%, recovering 78.03% of the decision gap between the base-state and reference-state conditions. Error-level analysis further shows that different state-error categories are associated with markedly different downstream decision gaps. These results show that making the right user evidence available is not sufficient for reliable personalization: models must also correctly interpret what that evidence implies about the user. Our findings motivate treating Evidence-to-State reasoning as an explicit evaluation target in personalized LLM systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.