Guided Reasoning Across ISA Dialects: Inference-Time Control for GPU Assembly Translation
Abstract
GPU software is increasingly shipped as compiled NVIDIA binaries, while the hardware that runs it is diversifying. Porting such software to another vendor usually requires source code that is often unavailable, or a hand-written binary translator for each instruction set. Large reasoning models can translate native GPU assembly without task-specific training, but not reliably: programs are long, reasoning and translation share one output budget, and the compiler can confirm that a translation is legal but not that it is correct. We treat translation as inference-time control of a stochastic model and present GRIT, which translates the native assembly in an NVIDIA executable into an AMD executable. GRIT splits programs into kernel translation units, grounds the model in context recovered from the binary, repairs rejected candidates with compiler diagnostics, and uses a learned Reasoning Controller to choose effort, context, repair budget, and attempts per program. On 375 programs from three benchmark suites, GRIT raises Gemini 3.8 Flash from 4.0% for whole-program translation to 77.3% in a single attempt, and the controller reaches 88.0% on a held-out test set with repeated attempts. We also find that more reasoning is not free: high effort exhausts the response limit six times as often as medium.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.