A Text-Free Preschool Mathematics Benchmark: What Its Scores Do and Do Not Measure
Abstract
We present a procedurally generated benchmark for preschool mathematics — counting, comparing, adding and subtracting with counters, pattern continuation, conservation of number and equal sharing — designed for a cascaded agent of small models rather than a single large one. The benchmark is rendered entirely without text: identity, size, position, container contents, holding state, episode phase and counting boards are encoded as colour, geometry and layout, so that no task can be solved by reading. Where a family admits a non-verbal answer the agent must demonstrate the result by acting, and both cascade stages are educated on corpora generated from the same distribution, shipped with the benchmark. We evaluate 21 configurations over 608 evaluation episodes and find that the training method, not the architecture, accounts for the largest gap in the results: full-parameter training of a 2B front end reaches 0.809 where low-rank training of the same model reaches 0.480, and a quarter of the education data with full-parameter training beats all of it with low-rank training. The front end bounds overall performance, with the back end contributing a family-specific competence profile that does not scale monotonically. We then audit the benchmark itself, computing for each family the score attainable by an agent that neither perceives nor reasons. Sixteen of 19 families carry usable signal, two are saturated, and one has none: conservation of number reports a peak success rate of 0.53 against a floor of 0.500, so its apparent difficulty is an answer prior being collected, and no configuration clears its own constant strategy by more than one item. We argue that this floor computation belongs in the benchmark's release gate rather than in post-hoc analysis, and that for families with skewed answers the headline metric must be balanced accuracy reported against chance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.