TrainTailor:A Coding Agent for Synthesizing Model-Specific Training Engines
Abstract
Large language models (LLMs) are released at a rapid pace, requiring training frameworks to continually adapt to new architectures. However, adapting general-purpose training frameworks to these architectures requires substantial engineering effort, while their shared abstractions can introduce runtime overhead and constrain model-specific optimization. In this paper, we explore an alternative enabled by recent advances in coding agents: building a dedicated training engine from scratch for each model. We instantiate this idea as TrainTailor, an agent-driven system that synthesizes and optimizes a standalone training engine tailored to a given model, training task, and hardware environment. Specifically, TrainTailor first builds a functional training engine with an initial execution configuration for the target hardware and validates its training behavior against a trusted reference. It then iteratively refines the engine implementation and execution configuration together, guided by profiling and bottleneck analysis, to improve training efficiency. Freed from the need to support diverse architectures, the generated engine can eliminate unnecessary abstractions and enable model-specific optimizations, such as architecture-aware operator fusion. Experiments show that the engines generated by TrainTailor satisfy the prescribed numerical tolerances over 30 steps, while improving model FLOPs utilization (MFU) by approximately 30% on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.