acceptodds
Under review as a conference paper at ICLR 2027

Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling

Abstract

Large language models are increasingly deployed as complex agentic systems. While prior work has focused on model- and system-level scaling, algorithm- and task-level scaling remain largely unaddressed. At the algorithm level, test-time scaling enhances reasoning but introduces cross-path redundancy: branches expanded from the same state re-generate overlapping tokens, yet existing engines execute them as independent requests. At the task level, agents exhibit highly heterogeneous runtime characteristics, yet current inference engines schedule all requests uniformly without awareness of agent roles. We propose Hive, a multi-agent inference infrastructure that enables algorithm- and task-level scaling. Hive provides a description front-end that captures per-agent behavior and supports test-time scaling algorithms. Its back-end introduces Logits Cache for algorithm-level scaling, which reuses intermediate logits across redundant sampling paths to mitigate cross-path redundancy, and Agent-aware Resource Scheduling for task-level scaling, which prioritizes KV-cache retention according to agent contributions. Experiments show that Logits Cache achieves an average speedup of –, and Agent-aware Resource Scheduling reduces the hotspot miss rate by –.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.