acceptodds
Under review as a conference paper at ICLR 2027

OARS: On Post-Training for Cost-Efficient Heterogeneous Multi-Agent Systems

Abstract

Frontier large language models (LLMs) are powerful agentic problem solvers, but deploying them at every step of a long-horizon task incurs substantial inference cost. Such tasks involve many rounds of LLM reasoning and environment interaction, yet not every step requires the same level of model capability. We introduce OARS (Orchestrator-Aware Role Specialization), which reduces reliance on expensive large models by organizing specialized small models to perform delegated work. Specifically, a frozen large-model orchestrator decomposes and delegates tasks to small-model workers through a unified work-order protocol that governs information isolation and exchange. Workers return only concise, task-relevant findings rather than their full execution histories, thereby reducing the orchestrator's computational workload and context growth. We post-train small-model workers within the complete multi-agent system. A central challenge is the *free-rider problem*: workers can be rewarded for collective success regardless of their own contribution, as a strong orchestrator or other workers may compensate for unhelpful worker reports. We address this with contribution-aware training that uses worker-level progress rewards to distinguish individual contributions from collective success. Our hierarchical credit normalization further prevents episodes with more workers or longer trajectories from receiving disproportionate training weight. On BrowseComp-Plus, OARS—with 8B workers assisting 230–550B-parameter models—improves accuracy by 6.3 percentage points on average while being up to 10.6× more cost-efficient than standalone large-model agents. On DR-256, OARS improves accuracy by 19.9 percentage points on average while achieving up to a 4.2× reduction in cost per correct answer. OARS further extends to agentic coding, cutting cost per resolved task by 66.5% on SWE-bench Verified while maintaining accuracy. We will release the code upon publication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.