Anti-General Intelligence: Specialization via Focal Forgetting and Hyperplasticity
Abstract
On many ML tasks, frontier performance is achieved by pretrained LLMs that are subsequently specialized to a task by post-training on data from that task. Specialization is typically assumed to come at the cost of forgetting previous abilities, which is welcome in a specialist deployed for a single purpose. However, recent work has shown that post-trained models may retain information even without explicit anti-forgetting measures. Across a wide range of domains, we observe that small post-trained LLMs do indeed retain their pre-specialization performance on non-target tasks, most completely when specialized with RL. We analyze the characteristics of non-forgetting in a saturated specialist, finding that task performance is carried by sparse task-selective subnetworks and that non-target tasks hold capacity the specialist never reclaims. We then study an intervention for the focal forgetting of task information by *Low-Rank Ascent*, finding that forgetting alone does not reliably help the target on its own, and develop *AntiGen*, a method that improves models beyond saturation by selectively increasing model plasticity in the non-target representation subspaces that ascent identifies. Applied to small specialists, *AntiGen* outperforms direct specialization and contemporary plasticity baselines on math and commonsense reasoning tasks, generalizes better in-domain, and focuses the loss of log-probabilities in ascent-derived subspaces.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.