acceptodds
Under review as a conference paper at ICLR 2027

Loom: A GPU Compiler for Scalar Functions over Finite Index Maps

Abstract

AI for science applications often mix two kinds of computations: dense matrix operations that benefit greatly from GPU architecture, and application of small scientific formula on an irregular, indexed data structure, that is a poor candidate for the modern GPU if implemented naively. During the evaluation of the scientific formula, separate kernels for indexed read, arithmetic and reductions may cost more time moving the intermediate arrays than the actual compute. We introduce Loom, a GPU compiler for scalar functions over finite index maps. The user supplies a fixed-shape JAX function and a description of where each array element lives; Loom fuses the computation into GPU kernels specialized around one contiguous axis of independent lanes for vectorization. It accepts declared algebraic constants such as and rewrites polynomial arithmetic so that it can apply exact cancellations. Loom lowers programs, including those built by JAX automatic differentiation, through StableHLO, generates many PTX implementations and selects one based on timing. We compare Loom to existing domain-specific libraries: on NVIDIA B200 and A100 GPUs, benchmarked for two sparse equivariant operations that are widely used in machine-learned interatomic potentials, Loom-generated hot kernels are up to approximately 2× faster. By lowering the optimization to the compiler level, Loom trades a costly first call, paid once per program, for fast kernels specific to the shapes of a given problem. Loom shows that specialized compilers can complement performance-driven domain-specific AI for science libraries.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.