Can Adaptation Substitute for Model Scale? An Empirical Study of Language Model Capability Gaps
Abstract
Smaller language models offer a practical deployment option, and a range of adaptation methods can narrow their performance gaps relative to larger models. However, choosing a method for a particular workload requires more than knowing its best attainable score. We present an empirical study across 16 generalization and discovery benchmarks, 8B and 27B targets, and five strategies spanning prompting, harness optimization, and weight updates. We compare method outcomes and score margins, separating heterogeneous endpoint inventories from gains against corresponding unadapted controls. The evaluated methods exhibit both large performance differences and near ties, and their ordering can change with the target model. Behavioral analyses reveal benefits concentrated in particular task strata and changes accompanying reduced output failures. Cost decomposition relates method choice to quality requirements, adaptation expenditure, and reusable supervision. We propose an experience-guided mechanism that prioritizes feasible probes under a budget and selects by target validation. Replay of its validation gate on four additional benchmarks reveals both successful choices and ranking errors. Code and experimental results are available at https://anonymous.4open.science/r/adaptation-study-57D6/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.