Simulating Organizational Data Lineage for Realistic and Controllable Evaluation of Enterprise Agents
Abstract
Autonomous agents are increasingly being used to produce complex enterprise artifacts such as spreadsheets, documents, and presentations. However, generating these scenarios remains difficult since realistic benchmarks require substantial manual effort, while existing synthetic approaches offer limited control over the relevance, diversity, and complexity of generated tasks. Evaluating generated artifacts also remains a challenge since no single reference captures the requirements that are practically imposed by different organizational stakeholders. We introduce LineageBench, a synthetic framework that systematically evaluates agents in realistic and diverse enterprise states through multi-persona data lineage. Rather than simulating organizational activity forward toward a potentially misaligned emergent state, LineageBench begins with a high-level deliverable and recursively works backward to identify the upstream artifacts required to produce it, along with the organizational personas responsible for them. We instantiate LineageBench with 30 organizational graphs spanning multiple industries, grounded in the O*NET taxonomy and usage statistics. We also leverage the lineage structure to construct rubrics for evaluating each artifact by collating the expectations of its producer, data providers, and consumers. By varying the lineage distance between a target artifact and the ancestors supplied to an agent, we evaluate 12 frontier agents on more than 500 tasks with controlled compositional complexity. Results show that the strongest agents achieve nearly twice the score of the weakest, yet similarly capable agents can require up to twice as many tool calls. Performance also declines by 31% on average as lineage distance increases from one to five hops. Together, these findings expose novel gaps in the capability and efficiency of enterprise agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.