Fewer Bits, More Tries: Recovering Capability Under Memory Constraints
Abstract
Large language models are generally more capable than smaller ones, but their memory requirements can make local deployment difficult. We ask whether part of this capability gap can instead be recovered by running a smaller or more compressed model multiple times and selecting among its outputs using public or model-generated tests. Single-user local inference can leave hardware underutilized, allowing parallel candidate generation to trade additional test-time computation for accuracy at modest latency cost. Across four code-generation benchmarks, verified retry recovers substantial accuracy lost to model size or compression. A 7B Qwen model with eight parallel attempts essentially matches a 32B model on MBPP+ and exceeds it on the other three benchmarks, while using about less model memory, running up to faster, and using less energy. We observe the same tradeoff across model sizes, compression levels, NVIDIA GPUs, and Apple Silicon, although its systems benefit depends on hardware and recovery depth. These results show that model capacity can be traded for verified test-time computation when correct solutions remain reachable within a modest retry budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.