acceptodds
Under review as a conference paper at ICLR 2027

Adaptive Transfer for Cost-Aware Bayesian Optimization of LLM Agent Configurations

Abstract

LLM agents can now autonomously perform long-horizon development tasks. An agent’s configuration, including its reasoning effort, tool access, and resource allocation, affects its performance, runtime, and monetary cost. An optimal configuration needs to balance these outcomes, as greater resource use does not necessarily yield higher utility. A full grid search on each new task is not practical because every configuration requires a costly, time-consuming run. Bayesian optimization can improve sample efficiency in search, and transferring knowledge from earlier tasks can further reduce search cost. The challenge is to identify what transfers when outcomes change across tasks and few target evaluations are affordable. In this work, we propose adaptive-level transfer (ALT), a cost-aware Bayesian optimization method that uses target evidence to adapt what is transferred from earlier tasks. Through systematic evaluation of every configuration of one agent across several tasks, recording agent action trajectories and outcomes, we obtain complete landscapes for exact replay of configuration search. Our analysis shows that configuration effects on several agent behavior features are relatively stable across tasks, while their effects on utility shift. ALT transfers configuration-to-behavior patterns from source trajectories and learns their utility contributions from target runs. It combines this behavior-level hypothesis with an objective-level hypothesis that transfers source utility surfaces. The target evidence decides how much each contributes to the prediction. Its acquisition weighs expected improvement against predicted cost, with the cost discount learned from target evidence on whether resource-intensive behavior raises utility. In cross-task search under monetary budgets, ALT reduces mean regret area by 4–12% relative to our main transfer baseline, ScaML-GP with expected improvement across all four evaluated utilities, with the largest reductions on the performance- and time-focused utilities. When sources include other repetitions of the target task, ALT outperforms its behavior-only ablation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.