acceptodds
Under review as a conference paper at ICLR 2027

The Token Games: Evaluating Language Model Reasoning with Puzzle Duels

Abstract

Evaluating the reasoning capabilities of Large Language Models is increasingly challenging as models improve. Human curation of hard questions is highly expensive, especially in recent benchmarks using PhD-level domain knowledge to challenge the most capable models. Even then, there is always a concern about whether these questions test genuine reasoning or if similar problems have been seen during training. Here, we take inspiration from 16th-century mathematical duels to design The Token Games (TTG): an evaluation framework where models challenge each other by creating their own puzzles. We leverage the format of Programming Puzzles — given a function that returns a boolean, find inputs that make it return True — to flexibly represent problems and enable verifying solutions. Using results from pairwise duels, we then compute Elo ratings, allowing us to compare models relative to each other. We evaluate 18 models — frontier systems, low-reasoning-effort variants, and open-weights models — on TTG, and closely match the rankings from existing benchmarks such as Humanity's Last Exam, without involving any human effort in creating puzzles and at a total cost of about $2,400 — a small fraction of the cost of sourcing expert-authored questions for comparable benchmarks. We also find that creating good puzzles is still a highly challenging task for current models. Overall, our work suggests new paradigms for evaluating reasoning that avoid saturation by design, and that allow testing models for other skills like creativity and task creation alongside problem solving.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.