acceptodds
Under review as a conference paper at ICLR 2027

Triage-PiEvo: Evidence Admission for Principle-Evolving Scientific Agents

Abstract

Large language model (LLM)-based AI Scientists automate increasingly large parts of scientific discovery, but noisy observations make their revisions unreliable. Existing principle-evolving agents act on every anomaly, whether its source is a new mechanism or an evaluation artifact; invalid revisions then bias evolution. To address this, we formulate revision as evidence admission: anomalies steer search immediately, but only independently verified evidence may rewrite persistent principles, and contradicted principles are retracted. We realize this in TRIAGE-PIEVO, which pairs an anomaly-responsive search portfolio with a staged verification gate that controls what enters principle state. Across four scientific tasks with three backbones, TRIAGE-PIEVO attains the highest average solution quality, improving over the strongest baseline of each backbone by +7.5% ∼ 11.8%; on Qwen3-235B it matches that baseline’s full-budget quality with one-third of the evaluation budget. In constructed stress tests, it admits none of the corrupted, hidden-source, or exploit evidence that ungated variants admit, while missing no planted mechanism. Plugged unchanged into three external research harnesses, it cuts false evidence admissions by over 60% and raises verified scores by 31% ∼ 43%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.