EvidLoop: Staged Verification for Agentic Model Search in Aerodrome Weather Forecasting
Abstract
Language-model agents increasingly turn research hypotheses into executable model edits, but improvements observed during adaptive search may not persist across training seeds or evaluation conditions. We introduce EvidLoop, a framework that links each evaluated implementation to staged acceptance criteria and separates development feedback from final assessment. Candidates advance through single-seed screening, paired five-seed confirmation, and one-time assessment of frozen models on temporal and spatial holdouts, including a predefined weather-event subset termed HAZARD. On the AeroWF aerodrome-weather benchmark, three language models explore four intervention streams, yielding 70 scored attempts, 24 screening survivors, and 11 confirmed candidates. The confirmed edits reduce development RMSE by 1.22%–3.76% relative to the parent's mean development RMSE. Seven satisfy a joint mean-based acceptance rule, fixed before held-out evaluation, that requires temporal improvement and limits spatial and HAZARD-conditioned degradation. Five improve all three held-out means. Four of the eleven confirmed candidates improve temporal and spatial means but exceed the HAZARD margin, including the candidate with the largest temporal gain. Accepted edits span error-dependent reweighting, horizon-conditioned robust penalties, and adaptive residual scaling. These results identify useful forecasting modifications and show that cross-seed development gains and aggregate held-out improvements can coexist with conditional regressions. EvidLoop makes the resulting acceptance decisions explicit and traceable to the implementations evaluated.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.