acceptodds
Under review as a conference paper at ICLR 2027

Error-Propagation-Aware Routing with Adaptive Horizons for Multi-Model LLM Agent DAG Execution

Abstract

AI agents increasingly execute multi-step workflows whose subtasks vary in difficulty and are connected by complex dependencies. Recent LLM-routing approaches reduce cloud cost by selecting models at the subtask level. However, none of these approaches jointly considers error propagation across dependent subtasks, dynamically controls how long routing decisions remain in effect, and concurrently executes runnable subtasks, which form a frontier. We present PARA, an Error-Propagation-Aware Routing framework with Adaptive horizons for multi-model LLM agent DAG execution. PARA represents each request as a subtask DAG and concurrently executes the subtasks in each frontier. At each routing point, PARA jointly determines a model allocation scheme and an adaptive commitment horizon—the number of consecutive frontier waves for which the routing decision remains in effect before being reconsidered. This adaptive horizon balances routing overhead against the need to respond to evolving execution states and error propagation. Across seven benchmarks, PARA reduces cloud cost by 68.50% and end-to-end latency by 43.15% relative to all-cloud execution while remaining within 1.71–6.75 points of its task quality. Compared with the original HybridLLM and Router-R1 baselines and Minions, PARA improves average score by up to 15.02 points while reducing average latency and cloud cost by up to 46.06% and 41.51%, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.