acceptodds
Under review as a conference paper at ICLR 2027

LetsPlay: A Gameplay Trajectory Dataset for Visual Game Agents

Abstract

Playing games through screenshots requires agents to interpret visual observations, remember past events, and execute actions toward a task objective. Training agents to acquire these capabilities requires gameplay trajectories that capture both what the agent observes and how it interacts with the environment. However, existing game benchmarks often rely on textual state descriptions or game-specific actions, while many visual game environments provide no gameplay trajectories for training. To address this, we introduce LetsPlay, a gameplay trajectory dataset for training visual game agents. The dataset aligns screenshots with recorded device actions using a unified action schema and further annotates each action with its reasoning, immediate goal, and expected outcome. In total, LetsPlay covers 25 games, comprising 2,036 gameplay trajectories and 315K action steps. Building on LetsPlay, we develop PlayAgent, a multimodal agent for long-horizon gameplay. The agent conditions its decisions on recent actions and gameplay context. Extensive experiments demonstrate the usefulness of LetsPlay for training visual game agents and the effectiveness of PlayAgent in completing game tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.