SelfAI: A Self-Directed Multi-Agent Framework for Long-Horizon Scientific Discovery
Abstract
LLM agents increasingly explore scientific problems on their own, yet they are judged almost solely by the final result they report. That result cannot distinguish an efficient investigation from one that spent most of its budget on redundant trials. We introduce SelfAI, a self-directed multi-agent system that treats experimentation as trajectory-driven decision making. Given a research objective and experimental scope, SelfAI formalizes the task, reasons over accumulated trajectories to plan subsequent trials, and decides when further trials are no longer worthwhile. To evaluate these decisions directly, we build a controlled benchmark in which all methods face the same executed experiments and observe each outcome only after selecting it. Across tasks spanning scientific computing, machine learning, computer vision, medical image analysis, and drug discovery, SelfAI finds the best configurations nearly as reliably as full-budget classical optimization methods and LLM baselines while using substantially fewer trials, whereas early-stopping baselines with similar trial counts miss them more often. Ablations show that trajectory analysis, planning, and stopping make complementary contributions. The evaluation further reveals that methods with similar final quality can differ substantially in how many trials they spend and when they stop. Code and VSCode extensions will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.