acceptodds
Under review as a conference paper at ICLR 2027

Humanize MLE Agents: Multi-Agent Session Alternation Beyond Single-Agent Limits

Abstract

Existing LLM-based agent approaches for long-horizon machine learning engineering (MLE) often use context engineering, implemented through custom harnesses that manage tools, skills, and accumulated experience. However, we find that as single-session execution lengthens, the agent's reasoning intensity declines alongside slowing progress, summarized as **long-horizon deceleration**, and the benefits of these custom harnesses vanish as the underlying model improves, summarized as **harness staleness**. These findings suggest that their benefits may not persist across versions, which motivates a workflow-level approach in which agents build on one another's work across sessions. We model these alternative agent sessions as a *context-reset semi-Markov decision process* (**CR-SMDP**) and establish sufficient conditions for alternation to exceed single-agent reference limits, derive an explicit experiment-budget bound, and characterize when session caps preserve this advantage. Guided by this analysis, we introduce *Humanize MLE Agents* (**HMA**), which alternates two agents on a shared workspace, limits each session to at most experiments and resets the context between sessions, without additional task-specific harness or prior knowledge supplements. Across 75 Kaggle tasks in MLE-bench, **HMA** reaches a **state-of-the-art** any-medal rate of 78.2%, exceeding the strongest open-source baseline by points with a quarter of its time budget. Moreover, among six configurations spanning five heterogeneous agent pairs, HMA improves any-medal rate over the constituent baselines by points on average, demonstrating cross-pair generality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.