Scale Buys Imitation, Not the Dividend: Post-Imitation RL Adds a Size-Invariant Gain in an Imperfect-Information Game
Abstract
Does the gain from reinforcement learning on top of an imitation model grow with the model? The answer decides whether compute goes into a larger base or into the RL stage, and it is rarely measured, because the gain is small against the variance of the tasks where it matters. We call the gain one RL stage adds over its own clone the RL dividend and measure it on a controlled ladder: five transformer clones (8M to 508M parameters) of one 1.4-billion-decision expert corpus in a four-player imperfect-information game (Riichi mahjong), one belief-state-distillation PPO recipe held fixed, a rules-accurate simulator, replicated runs, and reading rules fixed in advance. The dividend does not scale. From 22.5M to 155M parameters it is +80 to +110 Elo at every rung (25 PPO updates; 30 at 155M, which reads +64 at its 25th), replicated by an independent second run at two rungs, and the three post-RL policies finish within four points of even against each other on the same tables. One RL stage on the 22.5M model (about five GPU-hours on one consumer card, an hour of it PPO) beats the 155M clone on 59% [56, 62] of shared tables; scaling the clone seven-fold and training it longer buys 55% [52, 58], and scaling it from 22.5M to 56M buys nothing (49%). PPO does the same thing at every size (entropy collapses from 1.4 to 0.5 within five updates, 16-23% of held-out human decisions move, two independent 508M runs agree on 97.3% of decisions). At the ends of the range the dividend is run-dependent: 0, +36 and +53 Elo at 8M; +45, +77 and +99 at 508M, where two of the runs differ by 58 Elo on the same tables and one keeps rising to +110 with fifteen further updates. In two small imperfect-information games with exact metrics the dividend is bounded by the teacher's headroom, nil against a near-Nash teacher, growing with capacity against a near-deterministic exploitable one and growing less against a noisy one, the regime a human teacher is in; a ladder of language models pretrained on one corpus (Pythia, 14M to 410M), cloned from one noisy teacher on a synthetic task, takes the same dividend at every size. We also report how we got the shape wrong first: three silent infrastructure defects (stale trajectories from 384 games in flight, a value-head gate dropped in a refactor, mislabelled discard tokens in the deployed encoder) sat behind a spurious inverted-U with every training diagnostic inside its healthy band; two reproduced it on the repaired stack, one had shifted every Elo of the earlier one, and null readings on a fixed instrument and a code audit caught them. Single-run RL scaling curves are not evidence. Scale buys imitation; it does not buy the dividend.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.