HASTE: Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering
Abstract
Agents evaluated on MLE-Bench solve every competition from scratch; whatever an agent learns while engineering one solution is discarded before the next task begins. We present HASTE, an ML-engineering agent that accumulates plain-text skills in a three-tier hierarchy of global, domain, and competition scope, promotes learnings upward between tasks through LLM-driven abstraction, and loads only the tiers relevant to the task at hand. Across three seeds on the 22 MLE-Bench Lite competitions, using Claude Sonnet 4.6 with a CLI backend under a 12-hour cap per attempt, HASTE reaches a medal rate of 62.1% ± 1.5 (SEM); a replication on Claude Opus 4.8 reaches 66.7% ± 1.5, with the same six competitions failing to medal under either model. A multi-pass campaign that keeps each competition's best graded attempt reaches 77.3%, in the leaderboard's top band on a mid-tier model at half the typical budget; because attempts are selected on test outcomes, this figure is an upper bound. The accumulation loop operates end to end: each seed grows a store of 121 to 143 skills from an empty start, the orchestrator promotes learnings across tiers, loaded skills are cited by name in refinement proposals, including 15 cross-competition citations in the logged runs, and revisiting competitions with a matured store recovers 7 of 12 first-attempt failures. A system-level ablation over 8 competitions and 3 seeds places the deployed configuration at 83.3% ± 4.2 against 62.5% with no skills and 50.0% ± 7.2 with an unscoped frozen store; the arms differ in harness budget as well as loading policy, so we read the ablation as comparing configurations rather than isolating loading. All code, the skill store, run logs, submissions, and grading artifacts are included in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.