acceptodds
Under review as a conference paper at ICLR 2027

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration

Abstract

Automating scientific discovery requires iterative research pipelines that generate testable hypotheses, recover from experimental failures, retain useful experience across research cycles, and maintain strict data provenance. Existing systems often struggle with these demands: hypothesis generation and evaluation are frequently conflated, failed experiments are repaired without revisiting the underlying research design, and reported numerical claims are rarely verified against execution records. We present AutoResearchClaw (ARC), an open-source 23-stage multi-agent research pipeline that turns a research question into executed experiments, per-hypothesis verdicts, and a manuscript. ARC combines structured multi-agent debate, self-healing execution with Pivot/Refine decisions, registry-based result verification, configurable human intervention, and cross-run evolution that converts failures and research decisions into reusable lessons. We evaluate ARC on ARC-Bench, a 55-topic benchmark spanning machine learning, high-energy physics, metabolic modeling, statistics, and quantum computing. On 25 machine-learning topics, ARC scores 0.596, compared with 0.511 for AIDE and 0.419 for AI Scientist v2. Cross-run self-reinforcement raises Overall performance from a mean fresh-run baseline of 0.564 to 0.827 after three persistent iterations. End-to-end studies further show that targeted intervention at high-leverage stages outperforms both full autonomy and step-by-step oversight, while component ablations show that debate improves quality, self-healing improves completion, and registry-grounded verification prevents unsupported numerical claims in audited gate-passing manuscripts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.