Can LLMs Serve as Reference Anchors for the German Kangaroo Competition?
Abstract
Benchmarks for mathematical reasoning rarely preserve the incentives, visual content, and long-running human reference data of real exams, especially outside English. We introduce Kangaroo, a longitudinal benchmark built from the German Mathematical Kangaroo archive from 1998–2025: 140 exams and 3,886 multiple-choice items across five grade groups, 1,746 of them multimodal, with official human results for 115 exams across 23 years. We evaluate four frontier vision–language models in a single-turn exam mode with explicit abstention under the contest's negative marking. GPT-5 is strongest at 88.34% accuracy, but every model loses between 23 and 43 percentage points on multimodal items. We then ask whether model scores can serve as reference anchors for human exam performance, that is, whether a model's score on an exam helps predict the mean score of the pupils who took it. Across the 115 exams, human and model scores are negatively correlated, and the correlation shrinks toward zero once we account for grade and the share of multimodal items. Calibrations fit on earlier years predict later years no better than the historical grade mean: 4.12 against 4.16 points of mean absolute error for an equal-weight ensemble. Current model scores therefore cannot serve as reference anchors for human exam performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.