acceptodds
Under review as a conference paper at ICLR 2027

TiaoQi: A Solved Game as Exact Ground Truth for Concept Discovery and Its Validators

Abstract

Concept-discovery methods return features with names, and the field validates a name against the model, its activations or human labels, never against whether the concept is true. A strongly solved game supplies that ground truth. Every position has an exact value, so a claimed concept can be checked against it, and so can the test that validated it. We introduce TiaoQi, a benchmark on Chinese checkers that, to our knowledge, is the first to check every concept claim about a network against exact truth, and the first to use the same truth to grade its own tests: how often they pass a random direction, and how often a passing feature passes again. On it, a standard sparse-autoencoder pipeline returns few but real causal handles on the network's moves, 39–62 per agent where noise would earn 0–3.8. The handles sit on the game's own folk concepts, the features found for a fact are needed for play as a set, and the count replicates across seeds, layers, positions, dictionaries and rule variants. The same truth also grades four of the field's validation practices. A shuffled-label control certifies every dictionary where a matched random slice does not, a value-readout pass rule admits 1.8% to 5.5% of noise, a single feature's label is one draw, and decodable and used come apart in both directions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.