acceptodds
Under review as a conference paper at ICLR 2027

From Task-Oriented Agents to Personalized Teammates: Understanding Continual Adaptation in Cooperative Games

Abstract

In cooperative games, task performance alone does not determine whether an AI teammate is a good fit for a player. Motivated by the data available to game developers, we study personalization from abundant gameplay logs and sparse, noisy post-match endorsements, without requiring the target player to participate in training. Gameplay logs describe how the player acts, whereas endorsements provide imperfect evidence of how they want a teammate to cooperate. This creates a continual adaptation challenge: preference learning and subsequent task optimization pursue different objectives and can interfere with one another. Through controlled experiments in Overcooked, we isolate the effects of preference initialization, training partners, and reward design. Our analysis shows that, in multi-agent on-policy learning, task optimization can lead to different behavioral adaptations under different training-partner distributions. The same task objective can improve a learned division of labor under one interaction condition, yet replace it under another. Guided by this analysis, we initialize the AI from endorsements, learn an interactive player proxy from gameplay logs, and continue task-reward RL with that proxy. This simple procedure improves both task performance and preference satisfaction in our benchmark without requiring preference reward.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.