acceptodds
Under review as a conference paper at ICLR 2027

TenureBench: Benchmarking lifelong agent capability over a decade of evolving workplaces

Abstract

LLM agents are increasingly expected to work for long periods and to improve from experience over their lifetime. In real deployments, however, an agent works on its own without learning whether its answers were right, while the records it relies on keep changing. We introduce TenureBench to measure the lifelong capability of agents, that is, how an agent's capability changes as it keeps working. TenureBench replays 13,407 official documents from ten years of finance, medicine, and law to one persistent agent per field in publication order, keeps old versions next to new ones, and never shows the agent the correct answers. On a 300-task sample of its 920 tasks, the best of five frontier models solves 52.7%, and every agent is weakest overall at tracking what changed between two dates. Instead of improving over its lifetime, each agent does worse on later tasks. A long working history makes an agent more efficient, halving its search, but it does not make the agent more accurate. Given the source documents and the grading criteria without their answers, the same agent solves 93%. The agent has the ability to solve these tasks, and what it lacks is the professional sense, built over a career, of what a complete answer must cover. The time-driven protocol of TenureBench extends to any evolving archive to measure the lifelong capability of agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.