acceptodds
Under review as a conference paper at ICLR 2027

Rethinking the Evaluation of AI Companions through the Lens of User Retention

Abstract

AI companions are currently often evaluated under an unrealistic assumption that users always stay there and interact with AI. This guarantee of user retention does not reflect how we actually use AI in the real world, as users' willingness to engage largely depends on AI behavior. For instance, an AI intrusively prying into users' privacy may drive users away, yet retention-guaranteed evaluation will not surface this failure. To address this evaluation-deployment gap, we formulate user retention as an evaluation constraint shaped by AI behavior, and introduce Retentopia, a retention-constrained simulacrum for evaluating AI companions. Taking information elicitation as an evaluation target, Retentopia highlights current LLMs' ineffectiveness and a substantial trade-off between retention and elicitative AI behavior. More importantly, our analysis shows that considering user retention largely changes evaluation results in both measured performance (up to a 70.8% drop) and ranking among companions (GPT-5.6 Luna rises from 5th to 2nd out of six evaluated LLMs). These show that common retention-granted evaluation can provide misleading estimates of AI companions' performance in deployment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.