acceptodds
Under review as a conference paper at ICLR 2027

Fewer Bits, More Tries: Recovering Capability Under Memory Constraints

Abstract

Large language models are generally more capable than smaller ones, but their memory requirements can make local deployment difficult. We ask whether part of this capability gap can instead be recovered by running a smaller or more compressed model multiple times and selecting among its outputs using public or model-generated tests. Single-user local inference can leave hardware underutilized, allowing parallel candidate generation to trade additional test-time computation for accuracy at modest latency cost. Across four code-generation benchmarks, verified retry recovers substantial accuracy lost to model size or compression. A 7B Qwen model with eight parallel attempts essentially matches a 32B model on MBPP+ and exceeds it on the other three benchmarks, while using about less model memory, running up to faster, and using less energy. We observe the same tradeoff across model sizes, compression levels, NVIDIA GPUs, and Apple Silicon, although its systems benefit depends on hardware and recovery depth. These results show that model capacity can be traded for verified test-time computation when correct solutions remain reachable within a modest retry budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.