XStack: A Benchmark for ML Repo Migration
Abstract
ML code should run where researchers choose, not where it was first written. Today, machine learning code mainly lives in two ecosystems, PyTorch on NVIDIA GPUs and JAX on Google TPUs. Migrating repositories between them means building and testing across frameworks and hardware. This migration is (1) a practical need, since it unlocks different hardware for researchers, and (2) a natural benchmark: it involves different frameworks and hardware, covers whole repositories, and is long-horizon and challenging. We introduce **XStack**, a benchmark that evaluates agents on migrating ML repositories between these two ecosystems. The agent receives a PyTorch repository on GPUs and must deliver a JAX codebase on TPUs. Its output passes if it meets our grading criteria, based on the source repository's published result, authenticated and calibrated by us. Migration by agents remains unreliable: Claude Opus 5 passes 81%, GPT-5.6-sol 44%, and every other agent at most 38%. XStack is a first step toward freeing ML code from its ecosystem. By releasing it, we make progress toward that goal measurable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.