acceptodds
Under review as a conference paper at ICLR 2027

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Abstract

Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orchestrators, dedicated retrieval subagents, and more. While such harnesses expand, the use of more primitive but improved coding agents — where LLMs have direct access to the execution environment by using read, write, and bash primitives — have not received enough attention in the field. In this paper we find that, under equal time budget and frontier LLM backbone, open-source state-of-the-art harnesses provide no advantages over a single session of a minimal-harness coding agent baseline, pointing to the backbone as the primary driver for performance. Via a series of large-scale systematic ablation studies, we argue that the machinery layers become redundant in the coding agent setting. We conclude that the effort spent elaborating hand-crafted harnesses around strong models yields poor returns for current MLE benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.