When Is Component Sensitivity Compositional? A Causal Study of Neural Resource Allocation
Abstract
Neural resource-allocation methods commonly assign each component a scalar sensitivity and rank components independently. This raises two distinct questions: whether the score tracks singleton utility, and whether ranking by singleton utility remains useful under joint allocation. We separate these two questions and test them through direct interventions under a common, logged protocol across quantization, layer skipping, structured pruning, mixture-of-experts selection, and adapter placement. The resulting behavior forms a spectrum. Layer skipping exhibits small median pairwise interactions, yet one-shot singleton selection can still incur substantial held-out NLL loss at the allocation budget; pruning rankings depend on background sparsity; GPTQ-based group refinement and adapter placement show strong non-additivity; and expert interactions vary across the tested depths. Large additive prediction errors need not imply large selection gaps: the six-layer adapter placement realizes only 15% of its singleton-sum NLL prediction yet trails the tested last-6 heuristic by only 0.011 nats. By contrast, on Qwen2.5-14B one-shot singleton selection at a 25% layer-skip budget incurs 1.013 nats of held-out damage versus 0.540 for sequential re-measurement, and under sparse pruning backgrounds the dense ranking transfers worse than background re-measurement in 16 of 18 conditions—separating additive prediction failure from realized allocation loss. The 14B layer-skip NLL advantage does not yield a statistically resolved accuracy advantage on ARC-Challenge or HellaSwag. As an exploratory diagnostic, residual-stream perturbation alignment is positively associated with pairwise interaction magnitude in 11 of 12 runs. We conclude that component sensitivity need not behave as an intrinsic scalar: it is a configuration-dependent marginal effect in several tested regimes, and allocation methods should validate compositionality at the deployment granularity, background, budget, and metric.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.