Delegating AI Evaluation to AI-Generated Automata in Games
Abstract
AI benchmarks predominantly evaluate models based on their direct, interactive capabilities. However, a principled evaluation of AI usefulness should recognize that true intelligence externalizes cognition, creating durable artifacts that make future work reliable without continuous intervention. The central measure of an AI is shifting from what it can do through regular interaction in a live setting to the utility of the autonomous tools it can build to accomplish goals. We introduce Strautomata, a framework that evaluates models not as interactive players, but as builders of autonomous artifacts (executable policies or automata). We measure this artifact-generating utility within the space of well-known strategic board games. Across 20 distinct games, we sample 5,200 Python-based automata generated by 52 Large Language Models and execute over 2,500,000 offline, head-to-head matches to establish comprehensive Elo ratings. By separating code generation from execution, we ensure highly cost-efficient evaluation while testing vital LLM capabilities, placing our benchmark on the cost–alignment Pareto frontier.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.