acceptodds
Under review as a conference paper at ICLR 2027

Learning structured metacontrol

Abstract

Recent results in large language model reasoning underscore a well-known truth in classical planning machine learning: the ability to intelligently control how compute is allocated is paramount. For tree search, the mechanisms for making these decisions often rely on handcrafted rules, such as running a fixed amount of iterations of search before deciding on a move. This contrasts heavily with the evaluation component of search, which has benefitted enormously over the years from the adoption of learned, rather than predefined, evaluators. Here, we present a method where a controller learns to decide whether to act or keep planning (i.e., metacontrol) based on a learned embedding of the current planning tree. We evaluate our model's performance on chess play, showing that it is a strong meta-controller and beats existing baselines on efficient allocation of tree search iterations. We examine different ways of parametrizing our idea and show that a key component is the proper integration and propagation of estimator uncertainty in the evaluator. Finally, we interpret the learned halting mechanism itself.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.