acceptodds
Under review as a conference paper at ICLR 2027

TerraWeave: Grounding Policy-Skill Co-Evolution through Verified Execution for Earth Observation Agents

Abstract

Tool-augmented Earth observation (EO) agents solve multi-step geospatial tasks, yet successful tool execution does not guarantee that the dependencies required by downstream steps have been established. We introduce TerraWeave, which uses verified execution state as a shared reference for policy credit and skill evaluation. GeoState records observation-supported resources, relations, and readiness conditions. Grounded Dependency-Local Group Relative Policy Optimization (GDL-GRPO) assigns decision-local credit by normalizing verified progress among decisions sharing the same dependency opportunity. Skill candidates are compiled and execution-verified, then evaluated through frozen-policy ON/OFF continuations from the same executable snapshot to estimate incremental utility. Verified progress drives policy updates, while measured incremental utility selects skill versions for subsequent training. On the 1,169-task OpenEarthAgent test set, TerraWeave reaches 50.37/77.01 in native answer score (Ans.) and generation-tool execution (Gen.). It improves Ans./Gen. by 2.26/2.76 percentage points over a matched skill-augmented RL baseline and by 7.87/12.18 points over Standard GRPO. With the bank disabled, the co-evolved policy retains most gains over policy-only GDL-GRPO; enabling the bank adds runtime gains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.