Reading the Board, Playing Another One: Turn-Level Evaluation of Open LLM Agents on Textual Sokoban
Abstract
Aggregate outcomes and no-op rates obscure where an interactive agent fails. We report a turn-level evaluation of five open 3–8B LLM agents on textual Sokoban: 180 episodes and 1,561 steps across six conditions, where no level is solved and no agent docks a box. Two further experiments bound instrument failure and level difficulty. Replaying push-optimal solutions through the harness, repaired to log the solving step, solves all six corpus levels. On a fixed ladder of one to six optimal pushes, five models from three families solve 0 of the 60 episodes at or above four pushes; on the three lower rungs, success is slightly below a pre-registered uniform-random baseline at one push and above it at two and three. With one board per rung, this bounds the fixed ladder rather than establishing a general complexity threshold. We code every corpus turn against a ten-family behavioural taxonomy induced from three blind reading passes, with a contrast set recording locally correct elements of a turn. In 14 episodes of one model across the three reading passes and both temperatures, an accurate board table appears in the model's own output beside prose that reasons from a different board; the clearest instance recurs in all three executions. The move a thought describes and the action it emits also dissociate: a plain move described, a push emitted, or the reverse. On ladder rungs 1–3, a pre-registered blind control of 291 shuffled single turns finds no separation between solved and unsolved episodes on the pre-specified whole-turn contrast item. We therefore present the taxonomy as a description of observed behaviour in this setting, not as a validated discriminator of competent and incompetent turns. A further observation is carried as a metric argument, not a frequency claim: the most irreversibly lost levels are lost to legal, successful pushes, so a count of ineffective actions ranks the most damaging episodes as the best. Both annotators are language models, so we report kappa beside PABAK and prevalence and separately assess sampled label defensibility with 316 human verdicts across three instruments, never pooled. A third language-model annotator over the whole corpus has a substantially lower median in-band kappa with either of the first two (0.319 and 0.381) than they have with each other (0.576), each median over its own item band. Re-running an identical temperature-0 configuration three times, only 14 of 30 cells take the same path: fixed-seed greedy decoding did not reproduce trajectories in this serving configuration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.