From Task-Oriented Agents to Personalized Teammates: Understanding Continual Adaptation in Cooperative Games
Abstract
In cooperative games, task performance alone does not determine whether an AI teammate is a good fit for a player. Motivated by the data available to game developers, we study personalization from abundant gameplay logs and sparse, noisy post-match endorsements, without requiring the target player to participate in training. Gameplay logs describe how the player acts, whereas endorsements provide imperfect evidence of how they want a teammate to cooperate. This creates a continual adaptation challenge: preference learning and subsequent task optimization pursue different objectives and can interfere with one another. Through controlled experiments in Overcooked, we isolate the effects of preference initialization, training partners, and reward design. Our analysis shows that, in multi-agent on-policy learning, task optimization can lead to different behavioral adaptations under different training-partner distributions. The same task objective can improve a learned division of labor under one interaction condition, yet replace it under another. Guided by this analysis, we initialize the AI from endorsements, learn an interactive player proxy from gameplay logs, and continue task-reward RL with that proxy. This simple procedure improves both task performance and preference satisfaction in our benchmark without requiring preference reward.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.