acceptodds
Under review as a conference paper at ICLR 2027

MLToolBench: Learning Tool-Augmented Agents for Machine Learning Development

Abstract

Machine learning development agents must connect local evidence to code changes and verify that apparent improvements are valid. General code execution exposes evidence, but does not determine what to inspect or how observations should guide subsequent actions. We introduce toolbench, a diagnostic interface layer over MLAgentBench, and train policies to use it through guided demonstrations, supervised fine-tuning, and reinforcement learning. During RL, outcome-based GRPO optimizes task success, while SPICE uses verified solutions as training-only context to reweight supervision on sampled tool-call tokens. This signal measures solution-conditioned likelihood support rather than causal credit. We train on 80 tasks and evaluate on 25 held-out in-domain and 10 out-of-domain tasks. Diagnostic access alone is insufficient: adding toolbench reduces unguided Gemini 3.6 Flash from 66.8% to 57.2% in-domain, whereas tools with guidance reach 74.0%. With toolbench, the complete SFT-plus-RL pipeline improves Qwen3-8B from 24.0%/13.0% to 52.8%/39.0% ID/OOD success, and Qwen3.5-35B-A3B from 34.8%/31.0% to 64.4%/49.0%. These results show that diagnostic interfaces become useful when paired with training that connects observations to interventions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.