ARTISAN: Unlocking LLM Kernel Optimization on Unfamiliar Hardware
Abstract
Rapidly evolving AI accelerators open a recurring gap between what the hardware can deliver and what its software actually realizes. LLM agents offer a promising way to automate kernel optimization, yet they struggle on unfamiliar hardware. A central limitation is missing hardware-specific knowledge: what to optimize on the new platform, and how to realize that optimization effectively. We introduce Artisan, which learns this knowledge automatically from a handful of expert kernels. Rather than imitating expert code, Artisan contrasts expert and agent kernels to learn reusable diagnostic tools for finding optimization opportunities and guides for realizing them on the target hardware. A task's artifacts are retained only if fresh agents that never see the expert can use the resulting support to achieve a verified improvement, filtering out knowledge that fails to transfer. On 40 unseen AWS Trainium 2 kernels, Artisan achieves a geometric-mean speedup over the vendor-provided NKIPy baseline, with every kernel verified correct, versus for the strongest baseline that learns from the same experts. Adding Artisan to four existing agentic optimization pipelines brings their reported speedups from – with 24–40/40 coverage to – with 40/40.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.