RSI-Wild: Benchmarking AI Data Researchers with In-the-Wild Agentic Trajectories for Recursive Self-Improvement
Abstract
Recursive self-improvement (RSI) envisions AI systems that can help improve future generations of models. Recent benchmarks evaluate research agents on data-centric post-training tasks, but rarely expose them to the mixed-quality trajectory pools encountered in real-world data research. We introduce RSI-Wild, a benchmark for evaluating data research agents in this more realistic setting. RSI-Wild defines 23 error types across six categories and constructs execution-grounded failed trajectories, which are mixed with verified successes to form controlled mixed-quality pools. Our experiments reveal a utility–reliability gap. Frontier agents recover substantial value from noisy data, but the strongest reaches only 21.7% on the software engineering task, compared with 29% for human researchers. They fail to consistently produce better models across iterations, use research time inefficiently, and have difficulty translating external resources into concrete experiments. Error-level analysis further shows that agents more readily address errors with explicit contradictory evidence, while often overlooking locally coherent failures that require reasoning about missing actions or verification. Together, these findings suggest that current agents can make useful data improvements, but still struggle to turn feedback and external knowledge into sustained performance gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.