acceptodds
Under review as a conference paper at ICLR 2027

STAR: Trajectory-Aware Sentence Routing for Efficient Multi-LLM Ensemble

Abstract

Multi-LLM ensembles exploit the complementary capabilities of heterogeneous LLMs to improve inference without further increasing model scale. However, existing ensemble methods incur substantial computational waste by invoking multiple models at every generation step, while their reliance on local signals can result in suboptimal reasoning trajectories. In this work, we find that multi-LLM integration is sparse yet consequential at the sentence level: semantic disagreement arises only at a few critical states, while decisions at these states can shape subsequent reasoning trajectories beyond what local signals capture. To exploit this, we propose STAR, a sentence-level trajectory-aware routing framework that triggers multi-LLM integration only when needed and routes candidates according to their long-term trajectory value. STAR combines Sentence-level Disagreement Prediction (SDP) to predict cross-model semantic disagreement before candidate generation with Tree-Advantage Routing Optimization (TARO) to evaluate alternative trajectories and optimize routing using counterfactual advantages. Extensive experiments show that STAR achieves state-of-the-art accuracy among multi-LLM ensemble methods while reducing inference cost by up to 88.04%. Our code is available in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.