acceptodds
Under review as a conference paper at ICLR 2027

Solver.EXE: Learning to Solve Expert-level Problems via EXploration and Exploitation at Test-time

Abstract

An expert develops mastery not only from textbooks, but also through practice and refinement. We show that LLMs can do the same at test time: with the right self-evolution procedure, small- to medium-sized models can more fully realize their reasoning potential and grow into expert-level solvers. Motivated by the classical exploration-exploitation tradeoff, we introduce Solver.EXE, a multi-agent test-time learning framework built around a single frozen backbone and an online-evolving memory. Solver.EXE turns each reasoning problem into a miniature learning environment: a Reasoner exploits the current memory to produce rollouts, an Explorer expands the strategy space by surfacing new solutions, an Extractor distills insights and warnings from both kinds of rollouts, and a Curator consolidates them into a compact, evolving memory. Over several expert-level reasoning benchmarks spanning research-level mathematics and broader STEM, Solver.EXE consistently outperforms exploitation-only self-learning baselines and enables 30B-level open-source backbones (Gemma4-26B-A4B-it and Qwen3.5-35B-A3B) to match or surpass proprietary models, including Claude Opus-4.6 and GPT-5.2. Our analysis traces these gains to two latent capabilities of the backbone: bold strategy exploration and careful insight exploitation. These findings suggest that inference-time orchestration is a powerful and still largely untapped axis for advancing reasoning abilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.