AmBR: Agreement-Based Routing for Efficient Reasoning with Large Language Models
Abstract
Large language models demonstrate strong reasoning capabilities, but deploying them at scale involves a major cost-accuracy trade-off between inexpensive but weaker models and highly capable yet costly ones. We introduce Agreement-Based Routing (AmBR), a lightweight yet effective routing strategy that uses inter-model agreement as a proxy for answer correctness. AmBR first queries two inexpensive models and returns their shared prediction when they agree, escalating only disagreements to a larger (more expensive) fallback model, requiring no training, additional sampling, or model modifications. We evaluate AmBR across multiple reasoning benchmarks using a diverse set of open and closed models. Our results show that agreement between different models is a highly reliable and low-cost correctness signal, enabling substantial reductions in inference cost compared with strong baselines and prior works, while maintaining competitive or even better accuracy, demonstrating a practical and model-agnostic pathway to scalable reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.