UlumBench: A Benchmark for Evaluating Academic and Multi-Step Problem Solving in LLMs
Abstract
We present UlumBench, a benchmark for academic problem solving and controlled multi-step execution. Its 6,167 questions from 19 university courses include multiple-choice and open-ended tasks with text-only and image-supported inputs. A separate diagnostic collection provides 500 source-linked constructed problems with two to eight specified operations and intermediate references. Across five model families, the strongest academic baseline reaches 72.04% accuracy. On constructed problems, the terminal advantage of supplied correct states grows by 14–26 percentage points between two and eight steps. However, independent-task and whole-sequence controls do not establish a model-general growing dependency penalty. Verification reduces the measured penalty from supplied errors for four of five models under the recorded scorer, while cross-model plurality raises multiple-choice accuracy from 77.03% to 83.19%. These results motivate evaluating academic answers alongside intermediate-state use, sequence completeness, and error recovery, rather than interpreting a single final-answer score as a sufficient account of reasoning reliability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.