acceptodds
Under review as a conference paper at ICLR 2027

AscendInfraBench: Benchmarking Coding Agents for Full-Lifecycle LLM Infrastructure Engineering on Ascend NPUs

Abstract

LLM infrastructure must coordinate model computation, communication, and state management. We introduce a capability-oriented task construction method that specifies required work, permitted changes, and performance objectives while leaving implementation choices open. We use this method to construct AscendInfraBench, with 24 tasks spanning Optimization, Model Adaptation, and Open Build across training, reinforcement learning, and serving infrastructure on Ascend NPUs. We independently rebuild each submission, check numerical correctness and state handling, measure performance, and review the source against task requirements. State checks verify that training resumes from checkpoints, rollout uses updated policies, and serving preserves request-cache isolation. We evaluate 9 coding agents across 216 development runs, each with a 2-hour budget. GPT-6 Astra delivers valid implementations on all 24 tasks, including all 6 Model Adaptation and Open Build tasks, achieving a 1.63× task-weighted geometric-mean speedup over task baselines. Across agents, valid implementations meet the same task requirements through different choices in execution, model integration, and system organization. AscendInfraBench assesses agents’ ability to optimize existing code, integrate model support, and build complete LLM systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.