acceptodds
Under review as a conference paper at ICLR 2027

TileTune: An Efficient End-to-End Autotuning System for Tile-Based GPU Kernels

Abstract

Tile-based domain-specific languages simplify GPU kernel development, but achieving high performance still requires tuning parameters such as tile sizes and pipeline depths. Existing tuning approaches face a tradeoff between efficiency and generality: search-based methods are broadly applicable but incur compilation and benchmarking costs for each evaluated configuration, while heuristic-based methods reduce evaluation work but often require kernel-specific customization. Moreover, the explicit tile structure that this new generation of languages exposes remains largely unexploited, as existing tuners were designed for loop-level schedules or fixed kernel templates. We present TileTune, an end-to-end autotuning system that provides efficient, out-of-the-box tuning for tile-based GPU kernels, which balances generality, efficiency, and tile-awareness. At the algorithmic level, we propose a lightweight cost model using static program information rather than user-written, kernel-specific selection rules. It uses a shared analysis of explicit tile operations and resource requirements to rank and prune candidates while retaining empirically good selections. At the system level, we build a parallel task-processing runtime that reduces the cost of evaluating survivors through efficient dispatching, grouped compilation, and overlapping compilation and benchmarking. Across a range of kernels and hardware platforms, TileTune achieves end-to-end tuning speedup over the TileLang search-based baseline. Its shortlist retains the full-search optimum in all evaluated A100, H200 and B300 workloads.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.