SAR-Arena: Benchmarking Agent-Driven Research on AI Computing Systems and Architecture
Abstract
Research agents are starting to tackle systems and architecture problems as AI workloads and computing platforms evolve rapidly. Such work begins with workload-grounded bottleneck analysis, couples cross-stack mechanism design with engineering implementation, and requires long-horizon, measurement-driven optimization of quantitative system objectives; yet aggregate implementation scores reveal little about how agents make these decisions. We introduce SAR-Arena, a benchmark of ten recent, underexplored workload–platform problems for evaluating research in AI computing systems and architecture. Each task fixes its workload distribution, quality-constrained objective, and baseline while leaving the optimization mechanism open, and offers three expert-designed settings that progressively supply a bottleneck analysis or method proposal. SAR-Arena pairs rubric-based judgments of submitted diagnoses and methods with task-specific quantitative evaluators of runnable solutions, so intermediate reasoning and final system behavior can be examined together. We evaluate six model–harness configurations in 180 runs with eight-hour engineering budgets and extend 18 full-cycle trajectories to 24 engineering hours. The main experiment finds that strong diagnoses and proposals do not consistently yield baseline-beating implementations; relative to full-cycle research, the expert-proposal setting has higher mean engineering scores chiefly through gains on model–task pairs that already beat baseline in both settings. Longer runs yield task-dependent gains, and a nine-category trajectory audit documents recurring measurement, mechanism, and iteration failures, showing that SAR-Arena can reveal where plausible research plans stop short of validated improvement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.