ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step
Abstract
To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone, even in the absence of documentation. However, existing tool-use benchmarks expose semantic tool schemas in static environments, allowing agents to rely on prior knowledge rather than autonomous discovery. To address this limitation, we introduce ScrambleToolBench, an interactive terminal benchmark designed to isolate behavioral reasoning. By removing semantic cues and enforcing a continuous task curriculum, the benchmark requires agents to uncover hidden tool behaviors entirely through trial-and-error interaction. The benchmark further introduces dynamic challenges, including mapping drift, stochastic action failures, and temporal execution windows, to evaluate whether agents can revise and adapt their hypotheses as the environment changes. Our evaluation reveals that initial discovery does not translate into robust adaptation. Removing semantic priors reduces mean completion from 93% to 32%, which falls to 3% under combined dynamic stress. When mappings shift, agents exhibit belief inertia or search exhaustively rather than following the displacement pointers recorded in their own map. Increasing test-time reasoning improves completion rates and reduces redundant actions, but it relies on broad exploration rather than structured recovery. While equipping agents with persistent memory reduces compounding errors, they remain unable to efficiently infer structural changes, highlighting a gap in current agent reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.