acceptodds
Under review as a conference paper at ICLR 2027

hyperpath: Deep Learning Programs as Directed Hypergraphs

Abstract

Research code in machine learning is written twice. The first version is the idea, usually a diagram of values and the relationships between them. The second decides the loops over time, what to keep and for how long, how to batch, and where to compile. That second version takes the effort and carries the bugs, and when a language model writes it, someone still has to read it and check that it is right. We present hyperpath, a Python library in which a program is a directed (possibly cyclic) hypergraph. Each edge relates a set of values through an ordinary Python function and says nothing about when it runs. The user states which values they hold and which they want. The library compiles the one (acyclic) derivation the graph admits, naming what it could not reach where there is none and declining to choose where there are two. Users write time as a local relationship, x[t] and x[t+1], and the library derives each call's loop over time, how long it keeps each value and how much of the sequence runs at once, leaving the user to declare a loop's direction and the order of an update. We evaluate by reproduction, writing GPT-2 124M against llm.c's PyTorch trainer, ResNet-50 against torchvision's ImageNet recipe, DDPM against its CIFAR-10 reference and PPO and DQN against CleanRL. We re-run every reference on the same hardware, and each graph reproduces its reference's result at comparable cost. We then write nineteen questions and variants of three of these models and of a VAE, a GAN and nanoGPT. The diffusion model's reference writes its loop over the timesteps by hand four times, for sampling, for DDIM, for the images predicted along the chain and for the likelihood, where its graph writes no loop and derives all four.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.