acceptodds
Under review as a conference paper at ICLR 2027

OPTScientist: Toward Recursive Self-Improving Agents via Automated Optimizer Discovery

Abstract

Designing optimizers for modern deep learning remains a challenging scientific problem, requiring the joint consideration of optimization geometry, state dynamics, numerical stability, implementation constraints, and empirical generalization. Existing automated discovery methods typically search either over unconstrained code spaces or within narrowly parameterized optimizer families, trading flexibility for validity and interpretability. We introduce OPTScientist, a theory-guided multi-agent framework that casts optimizer discovery as a collaborative scientific process. Specialized agents propose hypotheses, instantiate candidate optimizers in an extensible typed DSL, and evaluate them through transformer pretraining, while a shared knowledge system accumulates candidate lineages, experimental outcomes, failure modes, and reusable lessons to guide subsequent search. The DSL provides compiler-checked validity while supporting conservative extensions when existing primitives limit promising hypotheses. Combining multi-fidelity evaluation, cross-scale validation, and controlled interventions, OPTScientist identifies proxy hacking and translates diagnosed failures into executable constraints. The system discovers eighteen optimizer designs that achieve lower target-scale validation bpb than all baselines in our main benchmark, with selected designs also improving downstream performance and transferring to vision after learning-rate retuning. These results position OPTScientist as a concrete step toward recursive self-improvement, where experimental evidence refines the knowledge, search space, and constraints that guide subsequent discovery.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.