acceptodds
Under review as a conference paper at ICLR 2027

SLM Reasons, LLM Advises: A Cost-Efficient Collaborative Reasoning Framework

Abstract

Deploying large language models (LLMs) for reasoning-intensive tasks presents a fundamental cost–accuracy trade-off: frontier models achieve strong performance but incur substantial inference costs, whereas small local models are inexpensive but considerably less reliable. We show that this trade-off can be substantially mitigated through selective collaboration. We introduce SRLA, a refinement framework in which a small language model (SLM) performs the primary reasoning and selectively seeks feedback from a LLM only when its reasoning is likely to contain an error. SRLA determines when and where external guidance is needed using Chunk-level Guided Scoring (CGS), a lightweight adapter trained with direct preference optimization on chunk-level correctness preferences. At inference time, CGS evaluates intermediate reasoning without access to reference answers and identifies the least reliable portion of the solution. Only when the resulting score triggers the refinement gate does SRLA invoke the LLM advisor. The advisor receives the original problem, the complete solver trajectory, and the identified problematic chunk, and returns a structured diagnosis with targeted corrective feedback. The SLM then uses this feedback to refine its solution, while otherwise reasoning entirely locally. We evaluate SRLA on five mathematical reasoning and six code generation benchmarks using Qwen3.5-9B as the local solver and GPT-5.4 as the advisor. SRLA matches or even exceeds GPT-5.4 used as a standalone solver on the majority of benchmarks while reducing per-problem API cost by up to 92% on mathematical reasoning and 94% on code generation. On several benchmarks, the advisor is invoked only rarely, yielding negligible external API cost while preserving strong accuracy. These results show that frontier LLMs need not perform the entire reasoning process: selectively using them to diagnose and advise a capable local solver can achieve frontier-level reasoning performance at substantially lower cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.