Hard, Not Rare: Human Item Difficulty, Not Concept Frequency, Orders Model Errors Across Professional Examinations
Abstract
A common explanation for language model errors on specialised tasks is that the relevant concepts are rare in pretraining data. We test this hypothesis on individual items of the Japanese National Medical Licensing Examination, pairing exact n-gram counts from the evaluated models' own pretraining corpus with per-item human examinee accuracy, and extend the analysis to Korean licensing examinations in law, psychology, management, and chemistry, for over 4,800 items in total. The hypothesis does not hold at the item level. On 1,213 single-answer medical items, concept frequency contributes 0.30 percentage points of cross-validated R² at 32B, while human item difficulty contributes 8.71 points. In a positive control, the same counts predict per-token log-probability at r ≈ 0.76 within fixed-length strata, so the negligible contribution on examination items cannot be attributed to measurement error. Difficulty and frequency are near-orthogonal, with a rank correlation of 0.058, and models fail hard items regardless of whether the concepts are rare. Scaling from 7B to 32B adds 25.5 percentage points on the easiest items and 6.2 on the hardest. Human item accuracy also predicts model confidence in five models from four additional families, with r between 0.223 and 0.338, and in all twenty domain–model combinations on the Korean examinations, where every 95% confidence interval excludes zero and concept frequency contributes under 2 R² percentage points per domain. On professional examinations, the performance gap on hard items reflects task demands that pretraining frequency does not explain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.