acceptodds
Under review as a conference paper at ICLR 2027

KernelShift: Benchmarking Kernel Optimization Across Evolving Hardware

Abstract

Efficient AI computation depends on hardware that *evolves over time*, shifting the strategies needed to implement performant code. For CPU kernels, a key evolving factor is the set of available *architectural extensions*, which expose specialized hardware instructions that change the techniques used to optimize kernels. We study how effectively coding agents use these new, potentially unfamiliar extensions to write fast code. We introduce **KernelShift**, a benchmark for evaluating AI coding agents on kernel optimization across evolving CPU architectures. We design an agent-assisted, human-verified pipeline for transforming code repositories into benchmark instances, together with infrastructure that lets agents iteratively optimize code using feedback from the target hardware. Using this pipeline, we construct 85 kernels from real-world, actively developed repositories (ncnn, llama.cpp, Arm SIMD Loops, and KleidiAI), including inference workloads derived from real model executions. KernelShift supports incorporating new kernels and hardware configurations as libraries and processors evolve. We evaluate agents on proprietary and open-weight models across progressively newer architectural extensions (Arm Neon, SVE, SVE2, and SME2), examining how performance changes as hardware evolves. On established ISAs, agents exploit residual optimization headroom, with the strongest system achieving a **2.74× geometric-mean speedup** over the ncnn production reference on SVE. In end-to-end inference of two LLMs, agent-written kernels achieve a **1.3× prefill speedup** over llama.cpp's fast path. However, current agents do not yet reliably adapt to new extensions that require unfamiliar programming models, such as SME2's matrix tiles.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.