The Declining Hazard Reported for AI Agents Is a Pooling Artefact
Abstract
Published fits to METR's time-horizon data put the Weibull shape k below one, so the failure rate falls as tasks get longer, and read it as agents getting better as a task goes on. That decline is an artefact of pooling unequally difficult tasks: a mixed batch of light bulbs shows a falling burnout rate while each wears out. Modelling task difficulty reverses the sign: the shape moves from 0.63 to 1.85, and 17 of 20 models exclude k ≤ 1, a count no null simulation reached. Across admissible ways of pooling the shape spans 1.56 to 1.85; the data reject the only one consistent with constant hazard. The sign holds on every clock; the magnitude does not: the same runs clocked by tokens, rated task length or elapsed time have median shapes 1.700, 2.045 and 2.625; the ordering repeats on a benchmark sharing no tasks, agents or harness. Humans on the same tasks and clock are indistinguishable from constant hazard; agents exceed them by Δk = 1.225 [0.825, 1.725]. A blind reimplementation from the specification reproduced both headline values. METR's pooled horizons stand. At fixed task difficulty an agent's 80%-reliability length is about half its 50% length; the published shape implies a sixth. Report the sign; report the magnitude as a range across clocks, or not at all.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.