CurriculumWorldBench: Benchmarking Autonomous Reconstruction of Executable University Curricula from the Web
Abstract
University curricula are distributed rule systems: requirements, prerequisites, policies, and exceptions are scattered across institutional websites and must be composed correctly to support downstream decisions. We introduce CURRICULUMWORLDBENCH, a benchmark for autonomous reconstruction of source- grounded, executable curricula from the web. It contains 735 curriculum worlds from 111 institutions and evaluates agents both on what they extract and on whether the recovered rules behave correctly and whether claims of completeness are justified. Across models, scaffolds, controlled website layouts, and frozen institutional sites, we find a persistent gap between finding courses and recon- structing the rules that govern them. Even when course recovery is high, 63,432 of 84,036 decidable obligations are omitted entirely. Stronger models achieve higher whole program behavioral agreement, yet substantial coverage failures remain. Providing agents with the correct evidence pages improves reconstruction when they cannot search for sources on their own, but offers little additional benefit once search is already available. Agents also often report that reconstruction is complete even when the resulting curriculum remains incomplete or behaviorally incorrect. Together, these findings characterize curriculum reconstruction as a distinct agent capability. Success requires more than retrieving relevant pages or extracting individual rules. An agent must identify the governing sources, integrate their requirements into a coherent executable model, and recognize when important parts of that model are still missing or unresolved.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.