acceptodds
Under review as a conference paper at ICLR 2027

Orchestration within LLMs via Discovering and Orchestrating Internal Operators

Abstract

Large Language Models (LLMs) are increasingly composed into agentic systems, where orchestration has become a key factor in system performance. Rather than relying on ever-larger models, agentic systems can solve complex tasks with models of bounded size, which makes orchestration a practical alternative to scaling. This motivates internalizing such orchestration within a single model, which holds the promise to improve both performance and parameter efficiency. To achieve it, we propose a framework for internally orchestrating LLMs, which performs orchestration inside a single LLM through two modules. The first module identifies patterns of coordinated activity among attention heads, which we call modes, and measures their causal effects on subsequent reasoning. We find that modes fall into three functionally distinct classes, corresponding to exploration, effort, and monitoring. Building on these findings, we introduce a Mode-Guided Orchestrator that uses mode-derived signals to determine when to intervene, adaptively invokes mode-specific interventions, and aggregates results through sequential consensus. Together, these modules dynamically construct query-specific computation graphs at inference time. Across five backbones and multiple benchmarks spanning mathematical, scientific, medical, and multi-hop reasoning, our method lies on the accuracy-compute Pareto frontier compared to the selected baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.