acceptodds
Under review as a conference paper at ICLR 2027

CoverWeave: Coverage Percolation Governs Compositional Depth Extrapolation under Verifiable Rewards

Abstract

Reinforcement learning with verifiable rewards (RLVR) is now the standard recipe for training models on checkable reasoning tasks, and the capability it is meant to buy is composition: a policy trained on shallow derivations should solve deeper ones it has never seen. Which property of the training set decides whether this happens is unknown, because existing compositional benchmarks fix combination structure and data volume together and cannot separate them. We introduce CoverWeave, a procedurally generated symbolic environment with an exact checker and a hindsight data engine in which every legal rollout ends in a legal goal, so verified training examples are produced by symbolic execution at zero model cost. In this engine, pair coverage, the fraction of atom pairs that co-occur in training, is an exactly controlled variable at fixed trajectory count and depth profile. A percolation account predicts that depth extrapolation is governed by the giant component of the atom co-occurrence graph, with success at depth . The prediction holds without a free parameter: on a coverage sweep, matches measured success at three test depths within points on average. Coverage dominates volume: one quarter of the data above threshold reaches average solve rate, while four times the data below threshold reaches , no better than the base amount. Binder cumulant crossings confirm that the realized co-occurrence graph has its structural threshold at the random-graph value for vocabulary sizes to , with a finite-size exponent close to the mean-field value . A threshold of the same shape appears with no weight updates when coverage is varied only in the prompt, and its location agrees across six frozen models to within . Training above threshold transfers to building-block-constrained retrosynthesis, raising restricted-catalogue route success from to .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.