acceptodds
Under review as a conference paper at ICLR 2027

How to Train Your Little Dragon: Teaching a Small Language Model to Construct and Refine GPU Optimization Scripts

Abstract

Compiler transformations interact, making GPU kernel optimization a search for combinations that preserve the computation while improving performance. We investigate how to train a small language model to generate the optimization: a correct, performance-improving script that coordinates interacting transformations in one complete proposal. We introduce HOW TO TRAIN YOUR LITTLE DRAGON, a training method that teaches a small language model both to construct complete optimization scripts and to improve existing ones. Rather than learn only to rank search candidates or generate replacement kernel code, our LITTLE DRAGON learns to generate correct scripts that transform the under-optimized kernel and improve its performance. We devise a compact, lossless representation whose deterministic decoding reconstructs the exact compiler trace, preserving its transformation parameters and dependencies. We use a compiler framework TVM MetaSchedule to build training data on 54,116 distinct workloads, and construct 108,594 examples in two training sets: 60,236 teach the model to generate an optimized script from an unoptimized kernel, while 48,358 teach it to produce a better script given the kernel and a somewhat optimized script. Supervised fine-tuning teaches the model to construct and improve optimization scripts from these search-derived examples. We then add reinforcement learning with a shaped reward that combines correctness with measured performance, awarding a latency-based bonus only to numerically correct kernels. This enables the model learn not only from existing optimizations, but from the correctness and performance of its generated scripts. We apply this training method to a model with 8 billion parameters. After supervised fine-tuning and reinforcement learning, the model achieves per-attempt correctness of 91.27% for construction and 92.01% for refinement, while reinforcement learning produces faster kernels than supervised training on 88 of 102 construction tasks and 145 of 170 refinement prompts. With only 8 billion parameters, our LITTLE DRAGON achieves higher per-attempt correctness and better kernel performance than six frontier LLMs—GPT-6 Astra, Opus 5, Kimi-K3, and GLM-5.3 among them, each with hundreds of billions to trillions of parameters—for which task-specific training is unavailable or expensive. These results suggest that HOW TO TRAIN YOUR LITTLE DRAGON enables a small language model to learn from compiler search and direct correctness and performance feedback, internalizing an optimization policy that constructs and improves complete scripts in one shot.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.