GeoSkill: Benchmarking and Evolving Domain Skills for Remote-Sensing Multimodal Agents
Abstract
Ultra-high-resolution (UHR) remote-sensing interpretation requires coordinating local visual inspection and broader spatial analysis under task-specific constraints. Tool-augmented agents have advanced these capabilities, but effective execution does not itself make the underlying procedures explicitly reusable. Domain skills can preserve this expertise, yet their benefits and limitations in UHR tasks remain unclear. We introduce GeoSkillBench, comprising 1,385 evaluation questions across 13 UHR tasks, fine-grained skill annotations, a structured skill library, and a shared geospatial tool environment. With models and base tools held fixed, we compare agents without skill documents, with expert-curated skills, and with self-generated plans. Expert skills improve accuracy across the evaluated models, while their advantage over self-generated plans varies by model. Paired trajectories reveal gaps in spatial verification and tool execution. Building on these findings, we introduce GeoSkillAgent, which consolidates procedural corrections from failed trajectories on a dedicated training set to iteratively refine an expert-curated SkillBank. Candidate updates and the merged SkillBank are screened for validation accuracy gains and overall stability before adoption. Model parameters remain fixed throughout, and the final SkillBank is frozen for testing. Our refinement experiments show that learning from failed trajectories yields larger gains than learning from successful trajectories alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.