TiaoQi: A Solved Game as Exact Ground Truth for Concept Discovery and Its Validators
Abstract
Concept-discovery methods return features with names, and the field validates a name against the model, its activations or human labels, never against whether the concept is true. A strongly solved game supplies that ground truth. Every position has an exact value, so a claimed concept can be checked against it, and so can the test that validated it. We introduce TiaoQi, a benchmark on Chinese checkers that, to our knowledge, is the first to check every concept claim about a network against exact truth, and the first to use the same truth to grade its own tests: how often they pass a random direction, and how often a passing feature passes again. On it, a standard sparse-autoencoder pipeline returns few but real causal handles on the network's moves, 39–62 per agent where noise would earn 0–3.8. The handles sit on the game's own folk concepts, the features found for a fact are needed for play as a set, and the count replicates across seeds, layers, positions, dictionaries and rule variants. The same truth also grades four of the field's validation practices. A shuffled-label control certifies every dictionary where a matched random slice does not, a value-readout pass rule admits 1.8% to 5.5% of noise, a single feature's label is one draw, and decodable and used come apart in both directions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.