ANDES: Agent-Native Data Evolving Synthesis via Feedback-Controlled Experience Acquisition
Abstract
AI agents are increasingly used to automate LLM post-training, yet acquiring targeted, high-quality training data remains a major bottleneck. Open-web curation requires long-horizon search, filtering, and balancing, while static synthesis pipelines cannot revise their acquisition distributions from intermediate outcomes. To address this gap, we formulate agentic data synthesis as feedback-controlled experience acquisition, in which observations from each synthesis round guide where subsequent supervision is acquired. We instantiate this formulation in Andes (Agent-Native Data Evolving Synthesis). Given a trainer-specified capability target, Andes maintains an acquisition distribution over a data-synthesis space modeled by a self-evolving World Tree. Diagnostics of effective quantity, topical allocation, and logical diversity then redirect subsequent acquisition toward under-covered capabilities while preserving broad contextual coverage. This synthesis-level feedback enables lightweight adaptation of the training signal before model-level training and evaluation. Across four base models under strict compute constraints, Andes substantially improves automated post-training, achieves performance comparable to frontier agents on PostTrainBench, and exhibits robust cross-task generalization. Our project is available at https://anonymous.4open.science/status/ANDES-4DAE.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.