TerraEvo: Skill-Guided Self-Evolution of Earth Observation Vision-Language Models
Abstract
Remote-sensing vision-language models (VLMs) remain heavily dependent on human-annotated supervision, constraining their scalability across increasingly diverse Earth observation (EO) imagery. Self-evolution offers a promising alternative, yet sustaining the reliability and diversity of self-generated supervision remains challenging. We introduce TerraEvo, a self-evolving framework that enables Proposer and Solver roles to co-evolve from unlabeled EO imagery without additional human annotations or an external judge. At its core, skill-guided Socratic dialogue transforms the Solver's answer disagreement into targeted follow-up questions. Quality supervision and adaptive rewards encourage informative, visually grounded questioning, while skill balancing and knowledge preservation support sustained learning across diverse EO tasks. A composite label-free learning signal further integrates answer consistency and reasoning quality to guide Solver improvement. Across five remote-sensing benchmarks, TerraEvo achieves the highest average accuracy among the evaluated general-purpose and domain-specific models, while ablation studies demonstrate the effectiveness of its key components. Our code, data, and models will be released publicly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.