acceptodds
Under review as a conference paper at ICLR 2027

AgentFloor: How Far Up the Tool-Use Ladder Can Small Open-Weight Models Go?

Abstract

Agentic systems make many small tool-using calls, but it is unclear which of them actually require a frontier model. We introduce AgentFloor, a controlled benchmark that organises tool-use tasks into a six-level ladder, from instruction following and single tool calls to branching, recovery, and long-horizon planning, while keeping the tool environment fixed. We evaluate 16 open-weight models alongside five frontier models from three vendors. Small open-weight models perform surprisingly well on the simplest levels: even a sub-billion-parameter model follows instructions reliably, but the gap widens as tasks require chaining, branching, and longer-horizon control. How large that gap is depends strongly on which frontier model serves as the reference: conclusions drawn against GPT-5 differ substantially from those drawn against the strongest Claude models. Models with similar overall accuracy also fail in very different ways, and simple interventions rarely transfer from one model to another. These results suggest that routing decisions should depend on the structure of the tool-use task, not only on model size or aggregate benchmark scores. We release the benchmark, the evaluation harness, and the full run corpus.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.