Decomposing Adaptive LLM Team Gains with a Frozen-Team Counterfactual
Abstract
Adaptive LLM teams combine two interventions: a fixed collaboration policy and a search procedure that adapts it. Standard evolved-versus-single comparisons conflate their contributions. We introduce a three-arm evaluation that measures Team Lift (frozen team minus single-agent policy) and Evolution Lift (evolved team minus the same frozen starting team) under a common allocated token budget. On a static SWE-derived patch-output proxy—scored from generated patch text without running repository tests—the evolved team exceeds the single agent by 6.875 percentage points across the observed runs. Arithmetically, this total consists of 5.350 points of Team Lift and 1.525 points of Evolution Lift; on GPQA, the corresponding decomposition is 3.0=3.8−0.8 points. Exploratory controls show that fixed policies with calibrated or deterministic selection can reach or exceed the evolved policy’s proxy score, while cold-start traces motivate broader proposal coverage. Rerun-aware and execution-based analyses distinguish uncertainty over search trajectories from validity of the evaluation endpoint. These results motivate a practical standard for collaboration search: report whether adaptation adds repeatable, task-valid value beyond a strong frozen policy instead of attributing the full team gain to evolution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.