Where Do Head-Level Adaptation Gains Come From? Design-Space Evidence
Abstract
Head-level adaptation has an attribution gap: complete pipelines are evaluated, but individual stages receive credit. Existing methods typically fine-tune random or all heads, or localize important heads and update them directly, bundling design choices whose contributions remain unresolved. Direct updating further assumes that diagnostic importance implies intervention utility, leaving a consequential choice implicit: *Bridging*, the mapping from localized heads to intervention targets. On Qwen2.5 with GSM8K, adding bridging shifts a localization-guided configuration from 5.9 percentage points below a random-head baseline to 4.8 points above it. To disentangle these contributions, we formulate an Evaluatology-based self-contained adaptation design system \(SCDS\). SCDS represents localization, bridging, and fine-tuning as attribute designed objects \(DOs\), their composition into adaptation paths as a structural DO, and model, dataset, and run-level conditions as confounding objects \(COs\). A blocked factorial study with matched interventions examines 14 localization rules, three bridging rules, and four fine-tuning rules across four model-dataset blocks, alongside random-head and all-head baselines. Across all four blocks, the structural path, bridging, and fine-tuning exhibit significant effects, whereas localization has a small, non-significant main effect. Of the two bridging rules, representation bridging significantly outperforms both Localize-FT and Random-FT in three of four blocks, while information-flow bridging shows no significant difference from the best baseline in the fourth. Bridging can therefore help, but only when the rule matches the block. Generalization tests of the same adapted models on two new datasets broadly corroborate the structural and bridging-rule findings. SCDS thus turns pipeline comparison into evidence of which design choices contribute to adaptation gains and how their effects depend on composition and evaluation conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.