In Search of Lost Time: Temporal Competence in Large Language Models
Abstract
Today's large language models (LLMs) increasingly interact with time: AI agents schedule tasks, coordinate with people and other AI systems, and even decide how long to think. Yet LLMs only ever encounter time secondhand—in descriptions rather than experience. Does this hinder their reasoning and decisions that involve time? Our paper traces this question from a model's knowledge to its actions through four successive evaluations. We find that models know the rules of time, which can be read, but struggle to predict durations, which must be lived; while they are highly accurate (91.9–95.8%) on a collection of questions built from the theorems of various time-related literatures, they underestimate their own completion times by up to 99.5% and fail to identify which tasks will take them longer to run. And while they are better at estimating durations of human tasks, their duration judgments magnify certain human biases by 1.3 simply based on how an event is worded. In practice, we find that some LLMs disregard time costs completely when answering user queries, ignoring a 1,200-fold increase in wait time entirely. Together, our experiments reveal new challenges posed to LLMs as they navigate a dimension they cannot directly perceive, and highlight how blind spots in LLMs can even be simple things that humans take for granted.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.