Scaling Test-Time Compute by Learning to Orchestrate Exploration
Abstract
Large language models (LLMs) are increasingly used to solve hard reasoning problems with large search spaces, where success can require sustained exploration over substantial test-time compute. We study how today's LLMs use this compute and find that chain-of-thought (CoT) can be contractive: as reasoning progresses, models devote increasingly less computation to novel directions and instead concentrate on approaches discovered earlier. Existing reasoning harnesses broaden exploration, but do not eliminate this contraction. We therefore study a harness in which an orchestrator explicitly decides which directions to explore and delegates them to parallel execution models. Current LLMs still struggle in this role, exhibiting redundant delegation, premature termination, and substantial problem-solving within the orchestrator itself. We hypothesize that this reflects a mismatch between standard reasoning training and the requirements of orchestration: models are trained to solve problems within their own reasoning trajectories, but receive little supervision for deciding where additional computation should be allocated. To address this, we introduce LoRE (Learning to Orchestrate Exploration), which first midtrains a model on synthetic exploration trajectories to establish orchestration priors, then applies multi-turn RL within the orchestrator-executor harness to optimize downstream task success. Applied to a 9B-parameter model, LoRE sustains broader exploration, scales effectively to millions of generated test-time tokens, and performs well on research-level benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.