acceptodds
Under review as a conference paper at ICLR 2027

CRISP: Conformal Routing and Interruption with Statistical Protection for LLM Cascades

Abstract

Reliable deployment of large language model (LLM) cascades requires deciding both which model to query and whether its realized output should be returned. We introduce CRISP (Conformal Routing and Interruption with Statistical Protection), which calibrates prompt- and output-side scores against labeled model failures. Under exchangeability, CRISP gives finite-sample, model- and task-conditional bounds on two events: an actually failing model entering the eligible set and an incorrect output passing the stopping gate. These are fixed-model failure-pass guarantees, not a bound of the same level on selective error or on the adaptively selected portfolio. Across seven open-weight models and four reasoning benchmarks, CRISP attains 81.5% accuracy and 99.5% coverage at a 0.0794 relative token-cost proxy on a fixed 200-example test split. Together, the formulation and evaluation show how failure-conditional calibration makes both routing and stopping decisions explicit and auditable in heterogeneous LLM cascades.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.