Distill Before You Measure: Dual Pruning for Evaluation-Efficient GPU Kernel Optimization
Abstract
GPU kernel optimization requires repeated compilation, correctness checking, and performance measurement. When candidate generation outpaces target-device evaluation, device feedback becomes the search bottleneck. We propose teacher-distilled dual pruning, a two-stage framework for optimization under limited evaluation budgets. Offline, a strong language model analyzes public implementations, patches, and failure cases, distilling optimization actions, applicability conditions, coupled changes, and source-level checks into linked knowledge records. Online, these records guide two pruning layers. Search-space pruning excludes routes incompatible with task and hardware conditions and guides search and generation within the retained space. Candidate-set pruning combines failure knowledge with tool results to inspect generated programs, block confirmed violations, and guide repair and selection before measurement. The framework requires no training or fine-tuning of the online model and supports open- and closed-source models. On 100 KernelBench Level 2 tasks on NVIDIA RTX 4090, it finds faster implementations on 77% of tasks with at most three candidate evaluations per task. The evaluated baselines first reach the same coverage with five to nine evaluations, corresponding to a 40.0%–66.7% reduction in candidate-evaluation budget. Experiments on NVIDIA H100 and with GLM-5.3-Flash further demonstrate applicability across hardware platforms and online models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.