Shared Accuracy–Length Patterns in Reasoning Compression
Abstract
Reinforcement learning with verifiable rewards enables strong reasoning, and reasoning-compression methods seek shorter correct responses through different training objectives and length controls. We place their independently trained policies in a common observable coordinate defined by dataset accuracy and mean response length. Across the observed task, model, and controller extensions, low-parameter functions predict held-out policy endpoints in this accuracy–length space. One shared endpoint direction learned from one group of prompts captures most transferable variation on a disjoint group. Its endpoint values are approximately monotone in realized response length with nonlinear spacing. Simple exponential and hyperbolic one-parameter shapes capture this spacing with little predictive loss and remain predictive after correctness is averaged over prompts. Together, these results connect the dataset-level accuracy–length structure of completed compression sweeps to compact transferable structure already shared across individual prompts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.