Gradient Skill Decomposition: Interpretable LLM Skill Vectors for Continual Learning
Abstract
Humans describe learning in terms of named and often compositional skills, whereas LLM training records model change through high-dimensional gradient updates. Without a correspondence between these views, a gradient does not directly reveal what an example teaches or which capabilities later training should preserve. We introduce Gradient Skill Decomposition (GSD), a framework for human–machine skill alignment in model gradient update space. GSD extracts compact layer-wise structure from per-example gradients, aligns this structure with multi-label skill annotations, and maps the recovered skill directions back to the parameter coordinates used for fine-tuning. The resulting shared representation supports both skill interpretation and skill-guided intervention. Across SmolLM3-3B, Qwen3.5-2B, and MiniCPM5-1B, GSD predicts human-defined skill sets with Jaccard scores of 64.2%–87.1%. During continual learning, we use the aligned directions as preservation constraints by projecting fine-tuning updates away from the selected skill subspace. Compared with standard LoRA, GSD-guided projection improves protected-task accuracy across 6 independently fine-tuned tasks and 3 models, with broadly comparable new continual-task performance. After 6 sequential tasks on MiniCPM5-1B, the projection improves protected-task accuracy by 12.5–18.0 percentage points over standard LoRA across two task orders. A targeted subset ablation provides initial evidence that selecting particular skill directions can concentrate protection on the corresponding behaviors. Across both independent and sequential continual-learning runs, once the skill basis is constructed, GSD processes the same number of fine-tuning tokens as standard LoRA, with only small increases in adaptation time and peak memory. Together, these results show that GSD connects an interpretable account of what examples teach to actionable control over continual learning, turning human-interpretable skill directions into a practical mechanism for reducing forgetting across successive tasks while maintaining new-task performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.