DataOS: A Recursive Self-Improvement System Driven by Environment Data
Abstract
Recursive self-improvement (RSI) seeks to make AI systems capable of revising the procedures through which they develop and improve their capabilities. Through successive experiments, these systems accumulate environment data (e.g., datasets and execution traces) that can inform subsequent improvements. Using these data effectively requires preserving the conditions behind past outcomes and identifying evidence and resources relevant to the evolving system. We introduce DataOS, a four-layer system for data-driven RSI that couples task execution and system evolution through continuously maintained environment data. DataOS preserves provenance and dependencies across versioned artifacts and execution records, enabling agents to use accumulated evidence and reusable resources to diagnose limitations and make targeted revisions to data, models, and the harness. We evaluate DataOS on PostTrainBench for autonomous post-training and Curation-Bench for multimodal data selection. Under matched researcher models, reasoning effort, and budgets, DataOS improves aggregate scores over native coding agents by up to 33.5% and 6.4%, respectively. Ablations support the contributions of structured history, reusable operators, and harness evolution, while transfer experiments show that accumulated research state can benefit subsequent runs even when target-model weights are reset.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.