Procedure or Format? Measuring What a Reasoning Trace Teaches with Format-Residualized Codelength
Abstract
A chain-of-thought trace can lower a model's loss on related problems for two different reasons: it can teach a procedure that transfers, or it can merely prime the surface format those problems share. A score based on that loss reduction cannot tell the two apart, and we prove that no function of it can. We therefore measure the reduction twice, once on sibling problems that need the trace's procedure and once on a matched null set in the same format that needs a different one, and take the difference. Format priming is shared by both sets and cancels; we call what remains the format-residualized differential codelength R. Under the standard mixture view of in-context learning, R is, up to a separation term, the expected log Bayes factor the trace induces between its own procedure and the null's. On 224 clusters built from GSM-Symbolic, R separates valid from wrong-method reasoning by 0.251 bits per token (p = 1e-5) with trace length and answer shape controlled, and where the raw reduction ranks a fluent trace that performs no computation above a valid one, R ranks them correctly (p = 9e-7). On 216 controlled clusters in three domains under five open models, R falls with corruption severity and predicts an independent model's out-of-distribution transfer beyond length, raw reduction, and correctness (partial rho = 0.27, p = 2e-15). The result is a label-free, forward-pass-only test of whether a trace teaches the procedure a problem family needs. It requires a scorer that has not already solved the task, and it measures in-context value, which does not carry over to selecting fine-tuning data. Code and data are available at https://anonymous.4open.science/r/format-residualized-codelength-8D79.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.