CourseProject-Bench: Evaluating AI Agents in Authentic Course Environments
Abstract
AI agents are increasingly evaluated on long-horizon tasks involving tools, files, and complex environments, yet existing benchmarks typically rely on agents' pretraining knowledge or provide the necessary domain knowledge directly. Therefore, we introduce CourseProject-Bench, a benchmark of 30 university course projects spanning five disciplinary domains, each paired with a reconstructed hierarchical course workspace. These projects require Active Knowledge Acquisition: agents must access relevant instructional materials from a broader course environment and understand the conceptual or procedural knowledge they contain before applying it to the project. To evaluate these open-ended tasks, we introduce Process Correctness, a five-stage framework covering Task Understanding, Knowledge Access, Conceptual Understanding, Knowledge Operationalization, and Artifact Production. Across eight frontier agent configurations, the strongest achieves only 60.7% Process Correctness. Stage-level analysis reveals that Active Knowledge Acquisition remains challenging, with conceptual understanding emerging as a major bottleneck. We further find that model capability has a substantially larger effect than harness choice, while performance also varies across disciplines and artifact demands. CourseProject-Bench provides a diagnostic testbed for studying agents that must learn from complex environments and apply acquired knowledge to long-horizon tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.