acceptodds
Under review as a conference paper at ICLR 2027

Same Score, Different State: Post-Training Moves Pre-Answer Recoverability under KV Compression

Abstract

Serving a reasoning model means shrinking its key-value (KV) cache while the model is still reasoning. Eviction and a token budget both act on the partial trace before the answer appears, and no benchmark grades that state. We measure it with an answer-marker-audited midpoint protocol: truncate a completed trajectory at 50% of its reasoning tokens, reject prefixes containing answer markers, compress the re-encoded midpoint KV cache, and elicit a fresh boxed answer. The protocol separates midpoint recoverability R, whether a forced answer elicited immediately from the audited pre-answer state is correct, from full-chain accuracy and from conditional eviction survival, with exact KV accounting. Our central finding is that post-training moves R by tens of points while paired full-chain accuracy shows no detectable change, in a recipe-dependent direction. Conditional on a correct source trajectory, the Base 1.5B model recovers 77% of midpoint answers on GSM8K. Sixteen public RL checkpoints post-trained from that base, fifteen by five length-shortening recipes, recover 15-60%, all below it, falling monotonically along each recipe's declared dose and tracking realized trace length (Spearman ρ = 0.96). The deficit is also present at 8B; on the audited trio it survives a change of elicitation wording, question-only re-solving and, in large part, an answer-leakage audit, and it persists when the model re-reasons from the compressed cache. Paired GRPO arms reproduce it in all twenty runs over three Qwen-1.5B and two Llama-8B seeds. By step 200, R falls to 35-42% at 1.5B under a length penalty and to 45-50% under a correctness-only control, and to 58.5-61.0% at 8B; accuracy is undetectably changed in every run, and the penalty decline is deeper in every 1.5B seed and absent at 8B. Correctness-only arms retrained at an 8,192-token cap with the registered fp32 weights train yet stay at Base level, 17-37 points above their 1,536-cap twins in all three seeds, so this manipulation supports training-time length pressure as the mediator. The direction depends on the recipe: a public RL-only 7B checkpoint recovers 96.1% against its registered base's 85.3% (+10.7 paired, p = 3.8 × 10^-5) at indistinguishable accuracy, and a 14B pair repeats it more weakly (+6.8, p = 6.6 × 10^-3). Eviction survival is separately model- and task-dependent: compressor rankings reverse across datasets. Under prespecified logistic predictors, no feature family predicts per-trajectory tolerance better than length does. A grouped Clopper-Pearson calibration nevertheless meets per-cell conditional-survival targets on held-out splits (up to 72% KV savings at the 75% retention target, 48% at 90%), while uniform ratios and cross-task transfer both miss it. Two checkpoints with no detectable benchmark-score difference can thus sit 40 points apart in R: same score, different state, in either direction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.