Hardware-in-the-Loop RL for OpenCL Optimization on Mobile GPUs
Abstract
As large language models (LLMs) make rapid strides in code generation, excitement is growing about their potential to optimize system kernels, but the field's datasets and evaluation infrastructure are primarily anchored in the NVIDIA ecosystem. OpenCL offers a common programming interface across vendors whose portable implementations can benefit from hardware-specific optimizations. However, learning these optimizations requires executable, correctness-verifiable workloads and reliable feedback from the target device. We investigate whether LLMs can use live execution feedback during RL training to optimize OpenCL kernels for resource-constrained mobile GPUs. We build and open-source two datasets: OpenCLKernelBench, which adapts KernelBench Level 1 with OpenCL reference implementations, and OpenCLOC, 620 curated kernels mined from GitHub. Both pair kernel implementations with test cases and executable host code. We also present a framework for hardware-in-the-loop reinforcement learning with verified rewards (HIL-RLVR) that includes a custom training environment, protocols for quantifying and reducing device noise, and novel reward hacking interventions. Finally, we train models to optimize kernels for Adreno GPUs on Snapdragon mobile processors. Relative to their shared Qwen2.5-Coder-32B-Instruct base model, increases from 9% to 26% on OpenCLOC and from effectively 0% to 18% on OpenCLKernelBench, achieving performance competitive with GPT-OSS-120B on both datasets. We also train an OpenCL policy on an NVIDIA Jetson AGX Thor and find that the learned optimizations are hardware-specific; policies remain correct when evaluated on the other device type, but their speedups do not transfer. To our knowledge, this is the first application of LLM-based RLVR to OpenCL kernel optimization. Our datasets and framework make resource-constrained deployment hardware a practical source of feedback for learning kernel optimizations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.