acceptodds
Under review as a conference paper at ICLR 2027

A Text-Free Preschool Mathematics Benchmark: What Its Scores Do and Do Not Measure

Abstract

We present a procedurally generated benchmark for preschool mathematics — counting, comparing, adding and subtracting with counters, pattern continuation, conservation of number and equal sharing — designed for a cascaded agent of small models rather than a single large one. The benchmark is rendered entirely without text: identity, size, position, container contents, holding state, episode phase and counting boards are encoded as colour, geometry and layout, so that no task can be solved by reading. Where a family admits a non-verbal answer the agent must demonstrate the result by acting, and both cascade stages are educated on corpora generated from the same distribution, shipped with the benchmark. We evaluate 21 configurations over 608 evaluation episodes and find that the training method, not the architecture, accounts for the largest gap in the results: full-parameter training of a 2B front end reaches 0.809 where low-rank training of the same model reaches 0.480, and a quarter of the education data with full-parameter training beats all of it with low-rank training. The front end bounds overall performance, with the back end contributing a family-specific competence profile that does not scale monotonically. We then audit the benchmark itself, computing for each family the score attainable by an agent that neither perceives nor reasons. Sixteen of 19 families carry usable signal, two are saturated, and one has none: conservation of number reports a peak success rate of 0.53 against a floor of 0.500, so its apparent difficulty is an answer prior being collected, and no configuration clears its own constant strategy by more than one item. We argue that this floor computation belongs in the benchmark's release gate rather than in post-hoc analysis, and that for families with skewed answers the headline metric must be balanced accuracy reported against chance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.