CreativeBench2.0: Benchmarking Creative Method Discovery and Foresight Toward Recursive Self-Improvement
Abstract
Major advances in machine learning often come from a small number of method changes that transfer across tasks and scale well. So far, these changes have been discovered by human researchers. Recursive self-improvement asks whether AI systems can discover them independently, but current evaluations cannot reliably measure this ability. They often allow broad changes to the surrounding training and evaluation pipeline, making it difficult to attribute gains to a specific component. They may also compare against reported results rather than baselines reproduced in the same evaluation harness. We introduce CreativeBench2.0, a repository-level benchmark for attributable method discovery. Each task is reverse-synthesized from a published research paper, its associated code, and its evaluation setup. The synthesis process isolates one replaceable algorithmic component while keeping the surrounding experiment fixed. A system may edit only this component. Its score is calibrated against reproduced baselines and evaluated on both visible and hidden out-of-distribution settings. The benchmark has two tracks. Discovery requires systems to implement and improve a method through experiments. Foresight requires them to rank methods and identify those that may fail under distribution shift, without executing code. Across seven systems and 50 tasks, reliable discovery is rare. Many attempts fail to produce valid experiments, and most valid attempts remain below the strongest reproduced baseline. Reference-context and interaction-budget comparisons show small observed gains, but runtime differences and uncertainty leave their isolated effects unresolved. Foresight systems rank methods above chance but remain unreliable at identifying distribution-sensitive methods. These results show that valid execution, cross-setting improvement, and pre-execution judgment remain distinct challenges.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.