acceptodds
Under review as a conference paper at ICLR 2027

What The Mask Still Buys: Measuring Which Decoding Constraints Small Language Models Internalize

Abstract

Schema-derived constraints let a small language model emit valid MLIR at inference time: a token-level grammar mask enforces syntax and type domains, and a scope validator rejects use-before-definition. We ask what the mask still buys once a model has been trained to satisfy it, layer by layer, and whether learning validity brings correctness with it. We train students on their own masked-and-verified rollouts (symbolic self-distillation, SSD) from 135M to 3B parameters in two families and measure, per constraint layer, the violation rate it still removes after training, paired per prompt on frozen pools. Validity moves into the weights: type domains internalize completely, syntax up to a structural residual of 5 to 11 points plus a vocabulary residual that shrinks when the training pool covers the missing concepts, and scope up to a small residual that a per-token scope signal mostly removes at 360M. Correctness follows only where the model already samples it. Verifier acceptance under free decoding rises by up to 70 points. Under a gold-differential harness, masked correctness falls at 135M, 360M and Gemma-1B; at 1.7B and 3B sampled pass@1 rises by 7 to 12 points, while the set of prompts solved within eight samples does not grow. We read both as concentration rather than acquisition. A functional reward does not change this on the released pools. Self-distillation raises the untrained model’s pass@1 and lowers its pass@16. Against the 49 of 143 prompts the untrained 360M model solves within 256 masked samples, nine on-policy cells each solve at most one prompt outside that set, while five cells that see a gold program solve 13 to 22, with no overlap between the groups. Supplying one gold program per rollout group restores correctness at a price in validity that an adapter control and a second seed show is not separable from adapter capacity at the scales we compare. As in reinforcement learning with verifiable rewards, self-distillation selects among programs the model already samples: it teaches validity, which the mask makes abundant, and not correctness, which the model must be shown; here the missing program can be supplied and its price measured. Outside the training distribution the student loses validity on held-out StableHLO above 360M while gaining it on a dialect that reuses the training vocabulary, and for Gemma-1B the mask does not repair the loss, leaving the student behind the model it started from with the full stack restored.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.