acceptodds
Under review as a conference paper at ICLR 2027

ForgeKernel: Self-Evolving Kernels Catching And Surpassing Experts And Agents

Abstract

Frontier coding agents now write GPU kernels that rival expert ones: left to run, even a generic agentic loop beats hand-tuned references on many operators. Yet every such loop stops at a ceiling of its own making. It settles a kernel’s architecture in its first few attempts and spends the rest of the run tuning beneath it, and each run starts from the same knowledge as the last and discards what it measured. We present ForgeKernel, a self-evolving kernel agent built to break that ceiling. It reads a reference kernel’s architecture from hardware behavior alone — running, timing and profiling a sealed binary, never its source — and turns it into an architecture-first curriculum that a coding agent implements and that is reopened whenever measurement shows the architecture, not the tuning, is the bottleneck. Every accepted kernel is sealed and promoted to the next campaign’s reference, so optimization continues past the best known implementation, and every finding certified by a hardware probe enters a knowledge base shared across operators. Across 27 operators on Hopper and Blackwell, ForgeKernel matches or beats 18 of 22 expert-library kernels, including FlashAttention-4, cuDNN and DeepGEMM, and all 7 published kernels of three kernel-generation systems, the latter by up to +170.6% (+39.4% in geometric mean). On DeepSeek-V4.1’s sparse-MLA prefill and fused attention kernels, released 15 days before this submission with no paper or second implementation, it comes within 1.3% and 3.7%. With only its profiler swapped, it leads the best public submission on all 8 level-4 CANN-Bench operators on Ascend, by +33.1% in geometric mean. Promoting the best known kernel to the reference, a generic loop’s included, gains up to a further +9.9%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.