VERA: Scaling Verifiable Envrioments for Agentic co-Evolution
Abstract
Competent agents need precise, verifiable environments: sandboxes that are resumable at any stage and evolve from observable evidence. Long-horizon work exposes how rare these are: a medical research agent must ground a finding, classify it, and write a report over dozens of dependent steps, yet recent environments score only the outcome. We present VERA, which builds such environments at scale and lets agents evolve on them. VERA builds these environments from initial trajectories: an agent writes rubrics and executable checks, a judge verifies each sandbox, and only those that pass enter the training bank. On these environments, VERA alternates between two updates: train the model with rubric rewards, or edit the harness skills; a verifier accepts each change only if it improves. This attribution distinguishes VERA’s co-evolution: updates target the cause, not just the outcome. We open-source 9,000+ long-horizon verifiable environments. At 9B, the co-evolved agent beats the strongest baseline by 10.3 and 13.0 points in the two domains. At 27B, it surpasses the baseline on AutoCoWorkBench (71.6) and AutoMedBench (80.7), transfers to unseen workflows, and retains general capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.