acceptodds
Under review as a conference paper at ICLR 2027

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Abstract

Video games provide a testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and action execution over multiple temporal horizons. Prior data and benchmarks either have narrow game coverage, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation suite measuring gameplay abilities at different horizons for diverse model families. GameHorizon Suite contains three components. First, GameHorizon-Annotator is a scalable and automated annotation pipeline for multi-horizon instructions. Second, using the pipeline, we construct GameHorizon-Data, a large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. It comprises 5,000 hours of recordings from 21 games. Third, we build GameHorizon-Bench with reproducible offline and stepwise online testing. The offline track enables reproducible evaluation via thousands of standardized questions in three primary tasks and ten diagnostic variants, while the online track tests whether offline scores reflect actual gameplay capabilities and localizes failures to specific steps within long-horizon gameplay. We evaluate 47 models through more than one million model invocations, showing a meaningful hierarchy of task difficulty and pronounced differences in model performance. Our work can provide a unified yardstick for gameplay abilities across models and horizons. We will release our data, annotator, and benchmark for future research.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.