ModelParts: Model-Grounded Tasks and Roofline Rewards for Training Kernel Agents on Emerging Accelerators
Abstract
As custom accelerators proliferate, kernel-optimization agents face a recurring cold start problem: the optimized libraries current methods use for training do not yet exist. Our premise is that the AI models a platform aims to train and serve can supply training data for reinforcement learning, while the accelerator supplies rewards based on its compute and memory throughput limits. We present ModelParts, which captures model computation graphs during compilation and explores alternative partitions to generate tasks whose kernel boundaries are not limited to existing libraries or manually selected motifs. To enable reinforcement learning on these tasks without baseline kernels in the target domain-specific language (DSL), we score measured latency against the hardware's roofline. We instantiate this approach with Neuron Kernel Interface (NKI), a domain-specific language for AWS Trainium and Inferentia accelerators. We construct ModelParts-6K, a corpus of 6,000 tasks generated from five models spanning distinct architectures. To isolate the effect of ModelParts data, we train models with reinforcement learning on ModelParts-6K and a synthetic dataset of PyTorch operator compositions. To evaluate transfer beyond generated tasks, we measure whether trained models can optimize the expert-written kernels in NKILib. With profiler-in-the-loop access, the ModelParts-trained policy yields a aggregate speedup on NKILib, surpassing both frontier baselines without profiler access.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.