E1: A First-Order Energy Metric for AI Model Architectures
Abstract
Energy is a primary constraint on AI inference at scale, yet models are designed and compared on accuracy alone. A trustworthy energy number requires an optimized implementation, which arrives long after the architecture is fixed, and only for the ideas that well-resourced teams choose to engineer. We introduce E1, an infrastructure-agnostic energy metric computed from unoptimized PyTorch code. E1 lowers the model to a graph of einsums, searches the tiling and fusion schedules of that graph on an abstract accelerator architecture with a finite on-chip buffer, and sums the energy of the arithmetic operations and off-chip data movement. The resulting E1 score represents what an ideally optimized implementation of the model could achieve, without any system optimization work by the developer. We validate E1 against energy measured on NVIDIA H100 GPUs for individual kernels, production LLM-serving kernels across implementations and across two years of their development, and for full model forward passes. E1 tracks measured energy over three orders of magnitude; optimized kernels land - above it and eager PyTorch about an order of magnitude higher; and when one mixture-of-experts model runs under two implementations whose measured energy differs by , E1 does not move. Our target is a reliable, infrastructure-agnostic metric that will enable fair accuracy-energy studies. We plan to open-source E1.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.