OpenMLE: Training LLMs for Evolutionary Search in Machine Learning Engineering
Abstract
Large language models have progressed from conversational assistants to reasoners and autonomous agents, and are now entering a stage in which they contribute to discovery itself, including the development of AI: AI that trains AI. Frontier labs have begun pursuing automated AI research, yet the open community lacks the training data and testbeds needed to study it: existing machine learning en- gineering (MLE) resources contain at most a few hundred tasks and are built for evaluation rather than training. We present OpenMLE, an open full-stack frame- work for AI-train-AI in executable MLE. First, OpenMLE-Gym provides, to our knowledge, the first large-scale AI-train-AI testbed: 5,758 quality-gated exe- cutable tasks spanning broad machine learning problems, with isolated execution, task-specific evaluators, and verified evolutionary trajectories. Second, OpenMLE-Train post-trains a single model for the discovery loop, where an agent repeatedly drafts, improves, debugs, and recombines solutions under delayed, noisy feedback: execution-verified SFT, then RL on per-task calibrated execution rewards, teaches it to serve as the shared program-revision operators. Third, OpenMLE-Evo deploys the trained model as the variation engine of long-horizon evolutionary search. On MLE-Bench Lite, post-training Qwen3.6-35B-A3B into OpenMLE-35B raises Medal Average from 39.39% to 60.61% with the harness fixed. With the model fixed, OpenMLE-EVO raises it from 53.03% under original AIRA-Evo to 60.61%. Adding benchmark-disjoint experience priors and asynchronous parallel search to OpenMLE-EVO, at the same total sandbox budget, brings the system to 71.21% with 3B active parameters, on par with GPT-5.5 with Codex (68.18%) under the same budget. Results on held-out NatureBench Lite show transfer to scientific reproduction tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.