acceptodds
Under review as a conference paper at ICLR 2027

Narrate to Simulate: Trajectory-Language Diffusion for Multi-Agent Scenes

Abstract

A simulated multi-agent scene is useful only if it can be read: not just where the agents go, but what happens between them. Generative models of multi-agent motion give the first and leave the second to the user. We extend joint continuous–discrete diffusion with a text modality: one denoiser generates the trajectories, the possession sequence and an event narration of a football play under one shared timestep, so every sample carries its own narration, grown from an empty sequence by an insertion-only edit flow. Insertion flows under-generate: a well-calibrated model reads its own short text as evidence of a short target and stops early, and more steps narrow the gap without closing it. , a one-line training augmentation, repairs it: at ten calls it improves narration content by a fifth and CIDEr two- to threefold, surpassing the unaugmented model at fifty calls and matching it at a hundred. We evaluate fidelity to the real play, coherence between the streams and control through the narration. On a season of a top European league and public World Cup 2022 data, the co-generated narration costs no significant accuracy against a matched baseline without it, only ball tempo, which tracks narration length; each narration describes its own play about as faithfully as the reference describes the real one; and editing it steers the play: swapping a pass target for a nearby teammate moves the ball to that player in most samples, with no retraining or guidance. Upon acceptance we release the league data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.