Adaptive Evolutionary Pareto Search with Online Transferability Estimation for Resource-Aware LLM Adaptation
Abstract
Adapting large language models to changing tasks and deployment requirements requires balancing accuracy, latency, and memory under limited evaluation resources. Although parameter-efficient fine-tuning (PEFT) reduces adaptation cost, selecting which layers to adapt together with rank and dropout creates a combinatorial multi-objective search problem. Under a fixed evaluation budget, reusing configurations from previous tasks can further waste evaluations when their usefulness on the new target is unknown. We present Adaptive Evolutionary Pareto Search with Online Transferability Estimation (AEPS-OT), a framework that allocates a fixed evaluation budget across multiple proposal mechanisms and historical configurations. A controller uses the current search state and objective weights to select among genetic algorithm (GA), differential evolution (DE), particle swarm optimization (PSO), and estimation of distribution algorithm (EDA) proposals, while a Pareto archive preserves search quality. To reuse experience from prior tasks, AEPS-OT first probes historical configurations on the target task and estimates transferability separately for accuracy, latency, and memory from rank agreement, stability, and performance gains measured on the target task. The resulting scores, combined using the current deployment preference, determine bounded source quotas, with all probes and transfers charged to the same 1,000-slot target budget. On Qwen2-0.5B and the Internet Movie Database (IMDb) dataset, five-seed independent audits give a mean hypervolume of 0.193424, higher than all fourteen external controls, including Bayesian optimization, fixed evolutionary operators, policy-learning baselines, classical multi-objective evolutionary algorithms and randomized search. In terms of mean audited hypervolume, AEPS-OT outperforms the base Adaptive Evolutionary Pareto Search (AEPS) by 4.4%. Search using target probes without subsequently adding historical configurations reaches 0.192410, indicating that evaluating historical configurations on the target task accounts for most of the observed transfer gain. These results support evaluating historical configurations on the current task before reusing them in resource-aware PEFT search under the evaluated single-model setting with one target task and four historical source tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.