acceptodds
Under review as a conference paper at ICLR 2027

Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards

Abstract

While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the resulting models remains poorly understood. In this work, we establish the first non-vacuous generalization bounds for parameter-efficient RLVR fine-tuning at the billion-parameter scale. Our approach adapts PAC-Bayes compression bounds to this setting and addresses the inherent stochasticity of token generation by applying the Gumbel-max reparameterization trick. To obtain non-vacuous bounds, we develop Progressive RLVR, a pipeline combining RLVR with on-policy distillation, TinyLoRA, and quantization, which achieves high empirical accuracy and high compressibility. Progressive RLVR retains 84–97% performance of standard LoRA fine-tuning while producing learned weight updates that are 14,796× more compressible. We show that this framework yields non-vacuous generalization bounds in four domains: math, programming, general-knowledge reasoning, and text-to-SQL. Our bounds exceed the accuracy of the base model by 9–51% and lie within 6–11% of the accuracy of the fine-tuned models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.