Logit Grafting: The Post-Training Delta is Sparse, Portable, and Powerful
Abstract
Post-training aligns a base language model, but the resulting behavioral change remains locked inside the post-trained model's weights. Decoding-time methods that take logit differences from auxiliary models and add it to a target model offer a viable solution to this transfer problem. We use logit grafting as an umbrella term to refer to this family of methods and use it as a a framework for studying how effectively this change can be extracted and transferred to a larger target model. This simple procedure has an elegant theoretical interpretation as an exponential tilting of the target's next token distribution. The resulting distribution uniquely optimizes a KL-regularized reward objective, giving the guidance strength a precise interpretation. Empirically, we find that the post-training delta is sparse, portable, and powerful. It only changes the sampled argmax token at 4–13% of the decoding steps and concentrates at high-entropy positions. It also transfers effectively across model families: a 1.5B Qwen delta, mapped through a vocabulary bridge, improves a non-specialized LLaMA-3-8B model's accuracy on the GSM8K math benchmark by 15%. Within the same model family, grafting a 1.5B delta onto a 7B base model closes 84–92% of the accuracy gap to the fully post-trained 7B model on math benchmarks, and the resulting model is preferred to the donor on most alignment and truthfulness comparisons. The grafted 7B model also outperforms the 1.5B post-trained donor in several settings without simply copying the donor's mistakes, suggesting a form of weak-to-strong generalization at inference time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.