AgriSEED: Learning to Teach Agricultural Vision-Language Models
Abstract
Generating useful supervision is a central challenge in adapting vision–language models to agricultural tasks. Although models can generate plausible question–answer pairs, answer validity alone does not indicate whether an example will help a solver learn. We introduce AGRISEED (Solver-Evaluated Example Design), a learning-to-teach framework that optimizes question generation using feedback from temporary solver adaptation. Given agricultural images, a questioner proposes question–answer pairs, and a frozen visual judge screens their quality. For each reliable candidate, a temporary solver undergoes a single low-rank adaptation step. Its learning gain is measured through changes in reference-answer likelihood on screened, self-generated probes from other images, complemented by augmented views of the candidate image. The temporary update is reset after evaluation, and the learning gain is combined with format and answer-quality rewards to train the questioner using GRPO. The trained questioner then generates a filtered dataset for a subsequent stage of solver reinforcement learning under the fixed visual judge. Experiments on AgroCoT with Qwen2.5-VL-7B, Qwen3-VL-8B, and InternVL2-8B show consistent improvements in semantic similarity and chainof-thought quality across all five agricultural task dimensions. A task-wise analysis shows consistent improvements in both semantic alignment and reasoning-step matching across the evaluated agricultural tasks and model backbones, supporting the effectiveness of AGRISEED for agricultural vision–language understanding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.