AIChipBench: State-Aware Evaluation of Coding Agents for AI Accelerators
Abstract
An RTL accelerator can emit the correct token while corrupting the state that later tokens depend on. For a coding agent, such delayed failures are hard to debug: the first visible mismatch may appear well after the faulty update. We introduce AIChipBench to evaluate whether agents can use public verification feedback to build implementations that pass independent acceptance as cross-module integration and persistent-state requirements increase. Each task pairs an executable fixed-point contract with a public validation ladder, an isolated held-out evaluator, and tool-use accounting, so that model fidelity, implementation correctness, and agent reliability are measured separately. On three controlled integration tasks with a four-layer Qwen-derived test model (Qwen3.5-nano), each attempted 15 times (five attempts from each of three agent configurations), final acceptance is 11/15 for a single layer, 7/15 for a four-layer one-step decoder, and 2/15 for eight-step recurrent decoding; all nine attempts that pass the public ladder but are rejected fail the held-out numerical checks. Scaffold-adaptation case studies cover Qwen3.5-0.8B and 4B; three Qwen reference designs satisfy their finite suites, while additional Gemma, Llama, and Qwen3.6-27B checks probe the limits of scaffold reuse. With arithmetic held fixed, paired schedules improve decode throughput by 12.5–15.0% at essentially unchanged LUT usage. The benchmark thus separates the feasibility of a hardware design from an agent's ability to construct it reliably, making state-preserving integration a concrete target for agent evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.