acceptodds
Under review as a conference paper at ICLR 2027

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Abstract

Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. That process is the training algorithm: a better objective or update rule improves the compute-capability exchange rate for every subsequent run, including the one that produces the next agent. Whether RSI is feasible therefore turns on whether an agent can design training algorithms. We present AI4AI-Bench to evaluate this capability on real machine-learning tasks, giving agents broad freedom to redesign algorithms and training frameworks within fixed task constraints. The benchmark comprises 10 frozen research repositories spanning 10 algorithm families. In each task, an agent has 4 hours on one B300 to rewrite the procedure; its submitted source is then rerun from scratch for up to 12 hours and scored by a fixed evaluator against the repository's reported reference. Because the 10 metrics are incommensurable, every task is mapped onto one scale on which is an uninformative or invalid result, is the reported repository reference, and is a fixed upper normalization anchor. Across 29 configurations of 6 systems on all 10 tasks the mean score is , and the best system reaches : even the strongest closes under a fifth of the reference-to-anchor interval. Among classifiable submissions, most match only run-level categories; patches matching at least one learning-level category average , compared with for run-only matches. Higher reasoning effort is associated with more learning-level matches ( versus among classifiable submissions) and higher scores ( versus ); the pooled effort groups differ in system composition, so these comparisons do not identify a causal mechanism. We release the task suite, the evaluators and every scored submission, so that the measurement can be repeated as these systems change.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.