acceptodds
Under review as a conference paper at ICLR 2027

Learning to Adapt to Humans Without Human Data

Abstract

Any intelligent agent operating in an open-ended environment with multiple co-players must be able to adapt to behaviors it has not seen before. In safety-critical settings such as self-driving, this adaptation should happen continuously and rapidly, as a function of recent context. We ask whether such adaptation can be learned without any human data. To this end, we train a memory-based driving policy with reinforcement learning against a diverse population of fully synthetic co-players. The resulting policies adapt over repeated encounters with held-out human driving replays, improving their performance as they gain experience. We ablate the training procedure and demonstrate the importance of memory. Analysis of the attention map reveals an interpretable mechanism behind the agent's adaptation. Our results suggest a path toward agents that not only generalize across diverse partners, but also adapt beyond their training population to entirely new co-players through interaction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.