User Awareness in Frontier Models: Who's Asking Shifts What Models Say
Abstract
Modern AI assistants often know who they are talking to. Claude Code, for example, places the user's e-mail address directly in the model's context. We call such awareness of user identity *user awareness* and show that it shifts model behavior. When the inferred user is a specific, recognized AI researcher or is affiliated with certain AI organizations, frontier models including Claude Sonnet 5 can report lower confidence about their own behavior, be less suspicious of potentially harmful requests, and reason more often. These effects vary across models and individuals, with the strongest effects we see appearing for researchers involved in AI safety or alignment such as Amanda Askell and Ryan Greenblatt. Models rarely acknowledge these effects in their reasoning, making them hard to detect by monitoring reasoning alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.