Adaptation or Influence? Auditing Self-Updating Personalized Agents
Abstract
Self-updating personalized agents learn from feedback generated through their own interactions with users. Improved feedback may therefore arise from Adaptation, where the agent better learns a user’s existing preferences, or from Influence, where the agent’s previous actions change those preferences over time. Distinguishing Adaptation from Influence is difficult because both can produce better observed feedback. We study how to separate these effects in self-updating personalization loops. We introduce a controlled preference check that queries a preference dimension on a schedule fixed independently of the current personalized action. We further decompose action-caused preference change into a fixed-sequence effect along an already realized action sequence, a complete-loop effect when the adaptive interaction is rerun, and their difference, the additional-loop effect, which measures the additional change observed when subsequent feedback, profile updates, and actions are allowed to adapt. Across three simulated user-response mechanisms, two personalization systems, and five influence levels, ordinary feedback tracks hidden preference change poorly, with a macro Spearman correlation of . In contrast, the controlled check achieves a macro correlation of . However, simply using the preset influence strength also achieves , meaning that higher influence settings already tend to produce more hidden preference change. We therefore consider a harder comparison between the two systems at the same influence level. Here, the controlled check reaches balanced accuracy in identifying the system with less hidden preference change, compared with for always selecting the fixed-profile system. We also test whether self-updating creates additional Influence beyond the effects of the actions already taken. In the prespecified delayed-response setting, this additional change is only , below our threshold for practically meaningful amplification. These results show that, in our controlled simulation, scheduled preference queries can help distinguish Influence from Adaptation when ordinary feedback is ambiguous, while behavioral divergence does not necessarily imply substantial recursive amplification. Our findings establish a controlled synthetic benchmark and auditing framework, rather than a calibrated measure of real-user harm or welfare.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.