Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards
Abstract
While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the resulting models remains poorly understood. In this work, we establish the first non-vacuous generalization bounds for parameter-efficient RLVR fine-tuning at the billion-parameter scale. Our approach adapts PAC-Bayes compression bounds to this setting and addresses the inherent stochasticity of token generation by applying the Gumbel-max reparameterization trick. To obtain non-vacuous bounds, we develop Progressive RLVR, a pipeline combining RLVR with on-policy distillation, TinyLoRA, and quantization, which achieves high empirical accuracy and high compressibility. Progressive RLVR retains 84–97% performance of standard LoRA fine-tuning while producing learned weight updates that are 14,796× more compressible. We show that this framework yields non-vacuous generalization bounds in four domains: math, programming, general-knowledge reasoning, and text-to-SQL. Our bounds exceed the accuracy of the base model by 9–51% and lie within 6–11% of the accuracy of the fine-tuned models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.