Human Difficulty Is Not Task Difficulty: What Policy Representations Know Beyond an Optimal Solver
Abstract
Human difficulty is not task difficulty. A policy network is trained to choose actions, not to judge how hard a problem is for people, yet part of it is linearly readable from its activations. Using chess as a model system, where difficulty is calibrated on large human populations, we probe eight networks. In Maia-2 a probe adds ΔR² = 0.0436 over gradient-boosted engine-search controls plus the model's own policy entropy, and 0.0350 with its own move predictions also controlled. Deeper engine search, at the depths tested, predicts human difficulty worse. The signal is shared across seven models after adjusting for the target, controls and a raw board encoding. We distinguish two measured aspects of difficulty: the population-calibrated rating and, separately, how long the original player deliberated, beyond that rating, the clock and engine features (the two readouts correlate at r = 0.59). Under matched ablations, neither direction has a distinctive effect on the solution move (the pre-specified primary outcome); the thinking-time direction shifts the move distribution (pre-specified secondary outcomes, small in absolute terms) without favouring or disfavouring the solution relative to matched controls. A person's error may also depend on how the situation arose: given the game so far, a human-imitation policy raises the probability of the wrong move a player made (0.111–0.136 nats, natural-log units, on months before Maia-3-79M's training window); an engine-trained network does not.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.