acceptodds
Under review as a conference paper at ICLR 2027

GeoToolAct: Evaluating Tool Execution in Remote Sensing Agents

Abstract

Remote sensing agents rely on multi-step tool calls to perform complex analytical tasks, yet reliable execution remains challenging. To analyze execution failures, we conceptually divide the model’s tool-call generation process into two stages: Tool Plan, which selects tools and their invocation order, and Tool Act, which constructs arguments after target-tool selection and before execution. Using this distinction, our trajectory analysis shows that Tool Act errors are less frequent than Tool Plan errors but more strongly associated with end-to-end task failure. This finding motivates examining argument construction alongside tool planning. However, existing remote sensing and geospatial agent benchmarks report argument accuracy but offer limited support for comparing performance across argument-construction requirements. To enable this comparison, we analyze argument-construction errors and identify three capability dimensions: adherence to the current tool schema (Schema-Prior), locating and faithfully reproducing required historical values (Contextual Value), and determining currently valid values after resource updates (State-Update). We build GeoToolAct around these dimensions by reconstructing observed failures into diagnostic tasks with fixed target tools and controlled non-target requirements. The benchmark covers 144 core tasks and 44 target tools, with four execution-history variants per task up to 120K tokens, preserving core tasks and reference answers. Evaluating 19 models on GeoToolAct, we find that the highest Overall Direct Accuracy is only 78.53% despite explicit target-tool specification, indicating substantial room to improve argument-construction reliability. Models also exhibit substantially different performance profiles, and sensitivity to longer execution histories varies across models and capabilities. Retries with evaluator-generated argument feedback yield uneven gains and incur additional generation costs. These results highlight the need for dedicated evaluation and targeted improvement of argument-construction capabilities to support reliable tool use in remote sensing.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.