Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling
Abstract
Large language models are increasingly deployed as complex agentic systems. While prior work has focused on model- and system-level scaling, algorithm- and task-level scaling remain largely unaddressed. At the algorithm level, test-time scaling enhances reasoning but introduces cross-path redundancy: branches expanded from the same state re-generate overlapping tokens, yet existing engines execute them as independent requests. At the task level, agents exhibit highly heterogeneous runtime characteristics, yet current inference engines schedule all requests uniformly without awareness of agent roles. We propose Hive, a multi-agent inference infrastructure that enables algorithm- and task-level scaling. Hive provides a description front-end that captures per-agent behavior and supports test-time scaling algorithms. Its back-end introduces Logits Cache for algorithm-level scaling, which reuses intermediate logits across redundant sampling paths to mitigate cross-path redundancy, and Agent-aware Resource Scheduling for task-level scaling, which prioritizes KV-cache retention according to agent contributions. Experiments show that Logits Cache achieves an average speedup of –, and Agent-aware Resource Scheduling reduces the hotspot miss rate by –.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.