Small Critics Make Strong Models Count: Learning to Verify Stronger Models for Efficient AutoResearch
Abstract
AutoResearch agents improve solutions through repeated experiments, but the intertwined demands of decision-making, execution, and analysis make them difficult to build and train. Using strong models throughout is also costly. We study which small-model capability yields the largest system-level gains from training. We introduce ResearchWeaver, which decomposes research into an Orchestrator, Worker, and Verifier. This makes research experience usable as role-specific supervision, so a small model can learn one responsibility rather than the entire process. We develop an automated task synthesis pipeline and collect role-structured trajectories on the resulting executable tasks. By training and replacing one role at a time while retaining strong models in the others, we identify where small-model learning helps most: a trained Verifier can influence subsequent research through its diagnoses and suggested next checks. Experiments on diverse benchmarks show that ResearchWeaver consistently improves performance across all three evaluated frontier models by explicitly structuring research into complementary roles. Training small models on role-structured experience further improves performance at the same source-task fraction, and role-specific experiments identify Verifier training as the highest-gain target. With supervision derived from only ∼2K source tasks, training small Verifiers delivers a 10.6-point gain on average, retaining 92.9% of the all-strong-model reference performance while reducing strong-model token use by 51.1% on average across various autoresearch benchmarks. These findings identify Verifier training as a high-leverage path to efficient AutoResearch.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.