Failure Modes in AI Retraining Dynamics
Abstract
Modern AI systems are increasingly retrained on data generated through interaction with users. Three forces are at play: (i) AI is typically retrained "greedily," ignoring exploration-exploitation tradeoffs, (ii) the users strategically adapt their behavior, and (iii) user intents are not directly observable. We ask whether these dynamics lead to poor outcomes. We study a stylized model, focusing on the "nice" case when the AI and the users have aligned incentives. We identify two distinct failure modes. First, the system may fail to converge to an optimal Nash equilibrium (of the relevant stage game) due to limited exploration, instead stabilizing at a suboptimal outcome region. This mode is ubiquitous: it happens with a positive probability for every problem instance. Second, a non-degenerate subset of problem instances exhibits model deterioration, whereby the system converges to an outcome that is strictly worse than the initial state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.