acceptodds
Under review as a conference paper at ICLR 2027

Learning When to Spawn Agents via Quality-Guided Cost Optimization

Abstract

A language-model recruitment policy must decide when additional agents justify their execution cost. With Group Relative Policy Optimization (GRPO), adding a cost penalty to the reward can make cost differences determine task-level advantages whenever all sampled trajectories have the same quality. We show why reducing the penalty coefficient may have little effect: normalization approximately cancels its scale when weighted cost dispersion is large relative to the numerical stability term. Quality-Guided Cost Optimization (QGC) uses the group's quality state to choose which cost comparisons to retain. It assigns zero task scores when quality is zero, compares team size alone when quality is identical and positive, and combines separately normalized quality and full estimated cost when quality varies. Positive quality includes partial success. The positive-quality branch therefore favors smaller teams at the same observed progress without penalizing output length. Across three Qwen3-8B training seeds on controlled compositional programming tasks, QGC raises mean success from 44.56% to 46.03% (+1.46 percentage points before rounding) and reduces mean estimated cost by 28.78% relative to Quality-Only. Against Guided Direct Cost, it gains 1.95 percentage points with 1.51% lower cost. The gains vary across seeds and models. Branch comparisons show higher success at slightly higher cost relative to full-cost tie feedback, and lower cost with a small success loss relative to tie suppression. These results support choosing cost comparisons according to observed task progress, while leaving matched-resource branch benefits to be established.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.