BharatMath: Culturally Regenerated Mathematical Reasoning for Twelve Indic Languages.
Abstract
Mathematics benchmarks for Indian languages are translations: they carry the dollars, miles and bake sales of their English sources into every language, they are small, and their English halves are saturated. We introduce BharatMath, 88,066 word problems in twelve Indic languages that are regenerated rather than translated. An executable formula is extracted from each English source problem, fresh operands are sampled and the gold answer is computed in code, so no language model can alter the stored contract; whether the generated question expresses it is audited separately. The question is then rebuilt around a record from a Context Bank: a finite, typed, enumerable relation of 782 curated cells that allocates regional units and scenarios (and, as advisory guidance, magnitudes) by language and mathematical domain, so that what the corpus is about is specified before generation and auditable after it. On identical contracts rated blind, the whole pipeline clears a cultural-plausibility criterion on 52% of items against 0.7% for literal translation; the bank's own contribution is +9 points of regional appropriateness, no significant change in cultural plausibility, and -5 points of mathematical equivalence. Golds are machine-checkable (99.94% reproduce from the stored formula); a two-family gold-blind LLM audit of test items finds 81.6% fully valid, 12.4% ambiguous, 3.7% with an altered operand and 2.1% with a differing answer, and the main effects below persist on the valid subset. Seven open-weight models span five points on MGSM-English but 2 to 86 points between their best and worst Indic language, and regenerated Bengali orders them differently from translated Bengali. Because every item is a contract plus a rendering, the benchmark also supports single-factor interventions, which locate six gaps: question comprehension (up to 52 points), numeral script (up to 34), a natural-language execution component for one model family, sampling-accessible headroom, the cost of forcing native-script thought under a fixed token budget (4-20 points, most of it truncation), and distractor robustness. We release the corpus, the bank, the lineage of every item, and keep a private held-out split.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.