acceptodds
Under review as a conference paper at ICLR 2027

Does Wider Search Help? Autonomous Research For Improving Vision-Language Models

Abstract

LLM-based research agents can autonomously propose hypotheses, modify training code, execute experiments, and substantially improve machine-learning models. Because model improvement is iterative, autonomous research can be viewed as search over research states and interventions, which raises the question of how much search control is actually needed. We study this question in autonomous post-training of a vision-language model by comparing a greedy single-trajectory policy () with a score-based multi-trajectory policy (). The two policies use the same research agent, starting model, training setup, evaluation pipeline, and research memory, and differ only in how many research states remain active. Both improve the aggregate score of the starting model from 0.576 to about 0.70, a relative gain of about 22% on average. However, under a round-aligned comparison that favors the wider policy, which trains three times as many candidates per round, reaches the plateau in fewer rounds but not a higher one: its best run (0.703) falls below the median greedy run (0.707). Analysis of the research histories shows that the wider policy retains diverse research states and attempts a broader range of intervention families, yet discovers no more productive ones, and that the effect of the same intervention can reverse sign as the surrounding training state changes, by margins well above training-seed noise. Our findings suggest that the large gains of autonomous post-training do not require wider search, and that exploiting additional preserved states, rather than generating them, is the central limitation for wider search in autonomous model research.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.