A SELF-GROWING EVALUATION METHODOLOGY FOR NPU OPERATORS WITH THEORY-DERIVED PERFOR- MANCE UPPER BOUNDS
Abstract
Deep neural network inference now spans edge devices and data centers, and its matrix- and vector-intensive computation calls for dedicated NPUs. Running a network on such an accelerator ultimately comes down to its low-level operators, whose quality directly determines model performance. Although operators can now be written in many ways, there remains no unified standard for whether an operator is correct and efficient. We therefore propose an NPU-oriented operator benchmark whose tasks are designed around the NPU's matrix/vector units and multi-level memory hierarchy, organized into four difficulty levels from vector-only operators to fused matrix-vector pipelines. However, fixed cases are insufficient for truly evaluating operators, as they capture only local behavior and can drive optimization into a local optimum. To address this, we propose SETG, which dynamically generates cases around two kinds of risk, namely operator numerical errors and NPU execution paths and overflow, while continuously refining them near anomalies based on real-device feedback. For evaluation, to determine how far each operator still is from its hardware limit, we define HAP, which analytically derives the performance ceiling from NPU scheduling characteristics, thereby measuring both achieved optimization and remaining headroom. Our experiments show that the benchmark both discovers real problems and aids effective optimization. On the one hand, in bare base-model evaluation, only four of eight base models obtain nonzero scores on all eight L1 operators, and the best average score is only 79.8, with many submissions failing on compilation or correctness, showing that the benchmark exposes real operator problems. On the other hand, in our internal evaluation, about 70% of operators score higher on hidden cases than on public cases, with average score up by 2.5 points and average speedup up by 0.51×, indicating that the optimization yields a genuine quality gain that generalizes to unseen cases, rather than overfitting to public cases.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.