acceptodds
Under review as a conference paper at ICLR 2027

TUNE-R2: Amortized Search Policies for LLM-Guided Hyperparameter Optimization

Abstract

Hyperparameter optimization (HPO) automates the selection of effective configurations in machine learning. Recent work uses large language models (LLMs) to guide this process through semantic priors and feedback from prior trials. However, most existing methods tightly couple reasoning with acting, repeatedly invoking the LLM over a growing trial history after each evaluation. This increases reasoning cost and can dilute decision-relevant signals with redundant context. In this paper, we introduce TUNE-R2, an amortized search policy framework that decouples LLM reasoning from individual objective evaluations. Instead of reasoning anew for every trial, TUNE-R2induces a reusable search policy from a bounded search state and refines it into a queue of candidate configurations for multiple subsequent evaluations. A state-adaptive policy horizon determines how many evaluations reuse each policy before it is refreshed with new feedback, balancing reasoning amortization against feedback freshness. Across OpenML, FCNet, and NASBench-201, TUNE-R2 achieves the best overall regret, AOC, mean rank, and win rate among evaluated LLM-guided HPO methods while requiring 68%–93% fewer Tokens@Best. Additional experiments on LLM fine-tuning and kernel-SVM HPO further support the generality of the approach. These results demonstrate the effectiveness of amortizing LLM reasoning through reusable search policies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.