acceptodds
Under review as a conference paper at ICLR 2027

From Post-Training Plasticity to Function-Targeted Adaptation in Large Language Models

Abstract

Post-training is central to developing and enhancing the capabilities of pretrained large language models (LLMs), yet its effects on their internal functional organization remain poorly understood. Inspired by experience-driven plasticity in biological intelligence, we investigate post-training plasticity: how the functional organization of attention heads evolves as LLMs acquire or enhance capabilities through post-training. We find that post-training primarily strengthens functional structures already present in pretrained models. This observation motivates Function-Targeted Adaptation (FTA), a selective fine-tuning paradigm that translates findings on post-training plasticity into targeted parameter updates. FTA identifies attention heads associated with a target capability and restricts updates to their corresponding weights, allowing us to examine whether existing functional structures provide a shared basis for both capability enhancement and knowledge suppression. We evaluate this hypothesis through skill learning and knowledge unlearning, assessing target adaptation alongside retention of non-target capabilities. Across multiple models and adaptation objectives, FTA updates at most around 10% of attention heads yet matches or exceeds full fine-tuning in target adaptation, while substantially reducing interference with non-target capabilities and retaining performance close to that of the original model. These results connect a mechanistic understanding of post-training plasticity with selective model adaptation, demonstrating how existing functional structures can guide both learning and unlearning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.