acceptodds
Under review as a conference paper at ICLR 2027

Can Smaller Models Help Larger Agents? Measuring the Incremental Value of Peer Advice

Abstract

Whether smaller models can provide more effective corrections than self-revision for the failed tasks of large language model agents remains a critical question, one that distinguishes a model's independent problem-solving capability from cross-model collaborative value. This paper conducts fair comparisons between pure self-revision and peer-assisted revision under unified model checkpoints and fixed revision budgets, evaluating all outcomes on hidden test sets across 300 competitive-programming tasks to quantify the gains and risks of peer advice. Advice from an 8B model raises a 35B MoE agent's success from 52.0% to 64.3%, and the gain survives assigning each packet to a different task; in the reverse direction, collaboration depends critically on the task–advice assignment, separating 63.7% under matched advice from 54.0% under mismatched advice. On shared tasks,the two matching effects differ by 15.2 percentage points, a difference reproduced across seeds. Yet a partner-free review that keeps the revision instruction but removes all advice leaves matched advice with no detectable advantage on a fixed 150-task subset, so part of the collaborative gain reflects a productive revision workflow rather than advice content; adding verified repair steps to the same diagnosis still improves the 35B agent by 6.4–14.0 points across three seeds.Accordingly, this paper establishes a paired evaluation framework that separates workflow gains, task-matching value, and content value,measures the increment of peer advice over partner-free review, and quantifies both its benefits and risks, offering new insights for lightweight collaborative optimization of large language model agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.