acceptodds
Under review as a conference paper at ICLR 2027

ORQA: Fast and Memory-Efficient Hybrid Reasoning with a Quantized Reasoner

Abstract

Reasoning improves the capability of large language models, but long reasoning traces increase latency and decoding cost. Hybrid reasoning reduces this cost by offloading intermediate reasoning to a lightweight model, which we call an efficient reasoner. Existing methods, however, require extra model parameters and KV cache, increasing memory consumption. Moreover, the capability gap between the reasoner and the target LLM often limits speedup or reduces accuracy. To this end, we propose ORQA, a hybrid reasoning framework that improves speed and accuracy over prior methods without adding memory overhead. Instead of loading an extra reasoner, ORQA uses a quantized counterpart of the target LLM and a shared KV cache pool. We propose weight infusion, which integrates parameters between the target LLM and reasoner, eliminating separate model storage while maintaining lossless inference for both. We further devise Quant-ifier, a routing signal designed for a quantized reasoner that predicts token disagreement using the quantization error distribution. Experiments show that ORQA improves speed and accuracy over prior methods with no additional memory.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.