acceptodds
Under review as a conference paper at ICLR 2027

TextCraft: A Text-Native Strategy Benchmark and Training Ground for Language Models

Abstract

Most language-model benchmarks are single-player: models are evaluated against fixed tasks that neither adapt as agents improve nor remain discriminative once performance saturates. Multi-agent competition offers a natural alternative, but existing environments are largely inherited from games designed for humans, where perception and interface complexity can obscure the language-model capabilities of interest. We introduce TextCraft, an open real-time strategy environment (tick-based, simultaneous-move) designed natively for language-model agents. Observations, communication, and actions are expressed entirely in text, while agents must simultaneously manage an economy, a technology tree, and military forces under partial observability across hundreds of sequential decisions. TextCraft supports both 1v1 competition and team-based 2v2 play, requiring long-horizon planning, decision-making under uncertainty, adaptation to an opponent, and coordination with a partner. Because difficulty is induced by other improving agents rather than a fixed answer key, TextCraft has no fixed saturation point and can remain discriminative as model capabilities advance. We release the environment, scripted baselines, and complete game trajectories from a range of open and closed models, enabling analysis of how language-model agents strategize, use partial information, adapt to opponents, and coordinate with teammates. TextCraft also serves as a testbed for improvement rather than measurement alone: With learned natural-language skills and no parameter updates, Claude Haiku 4.5 and Luna 5.6, two lower-ranked models on the leaderboard, achieve competitive win rates against stronger models at the final checkpoint under tournament configurations. Together, these results position TextCraft as a shared environment for measuring and improving long-horizon strategic and multi-agent behavior in language models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.