Accept, Repair, or Restart: Step-Level Risk-Aware SLM–LLM Collaboration
Abstract
Large language model (LLM)–small language model (SLM) collaboration offers an effective way to balance reasoning quality and computational efficiency by selectively invoking stronger models. However, existing methods typically determine collaboration through response-level confidence or fixed escalation rules, overlooking step-level failures in reasoning trajectories and the need for adaptive interventions tailored to different failure patterns. To address these problems, we propose step-level Risk-Aware Dynamic Action Allocation (RADA), a framework that formulates SLM–LLM collaboration as a risk-aware multi-action decision problem. RADA jointly models global response reliability and step-level reasoning risk to determine when LLM intervention is needed, adapts the intervention criterion using calibrated quality estimates of previously selected actions, and selects the intervention scope between local revision and full regeneration based on the identified failure pattern. Across mathematical and general reasoning benchmarks and multiple SLM–LLM pairs, RADA achieves the highest average accuracy among the evaluated collaboration methods. It improves average mathematical reasoning accuracy by up to 4.33 percentage points over the best-performing baseline while maintaining a favorable accuracy–computation trade-off.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.