Evaluating and Improving Human Behavior Modeling in Chess
Abstract
Human behavior models should reproduce not only likely actions, but also the stochastic structure and skill-dependent patterns of human decisions. We study rating-conditioned chess models and show that imitation policies can remain highly predictive while producing systematically mismatched move-quality distributions at higher ratings. To characterize this mismatch, we introduce three complementary diagnostics: mean highest-probability-density calibration error (mHPD-Err), which measures aggregate coverage calibration of high-probability move sets; move-class L1 error (Class L1), which measures differences in the frequencies of engine-derived move-quality classes; and move-class severity mean absolute error (Severity MAE), which measures differences in error severity within move-quality classes. We then propose an engine-informed information projection that minimally modifies the learned human policy to match population-level move-quality class frequencies and error severity. Across ratings 2100–2500, our method improves all three diagnostics relative to KL-regularized Monte Carlo tree search, while achieving comparable top-1 move matching accuracy and improving cross-entropy. At rating 2500, it reduces mHPD-Err from 5.10 to 1.73 percentage points, Class L1 from 1.89 to 0.43 percentage points, and Severity MAE from 1.027 to 0.010 centipawns, while also improving cross-entropy from 1.202 to 1.181 and top-1 move matching accuracy from 58.95% to 59.25% on the established ALLIE test set. These results demonstrate the importance of evaluating predictive accuracy, probability calibration, and skill-relative error structure as complementary aspects of human behavioral fidelity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.