acceptodds
Under review as a conference paper at ICLR 2027

ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents

Abstract

ScienceBuddy is an interactive scientific research workspace that brings evidence access, computational analysis, and researcher interaction into a single conversational workflow. At its core is recursive-in-recursive self-improvement, which couples harness evolution with model reinforcement learning: the inner recursion improves the harness with the model fixed, while the outer recursion trains the model under the improved harness. Harness evolution shapes training experience, and model learning creates new opportunities for harness adaptation. Given executable scientific tasks and independent verifiers, interaction feedback guides full-program revisions, while fresh verifier-scored rollouts train the model. We examine the resulting execution effects, training effects, and dependence on update order. Across seven scientific task categories, the reported two-cycle system reaches 70.4% Test accuracy from a 43.2% initial baseline. Crossed training and evaluation harnesses distinguish model learning from test-time program replacement; a separate control favors alternating updates over one-shot sequencing at equal proposal and optimizer-update counts. Reply-availability and editing-scope controls further characterize procedural adaptation. These findings connect the framework to observable improvements in scientific task execution, with task and verifier construction outside the evaluated method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.