Errors and path-dependencies in LLM descriptions of the world: An analysis via maps and routes
Abstract
When a language model describes a piece of the world in prose — a scene, a budget, a walk across a city — it produces a stream of small factual commitments: how many people were in the group, how much is left after the rent is paid, which side of the street the museum is on. We study two surprisingly robust properties of these commitments that are easy to overlook. The first is that models make discrete factual errors at a non-trivial rate even when the same model, asked directly, can identify the error. The second is that the content a model chooses to report is path-dependent: it changes systematically with the order in which material is presented or traversed, in ways that a faithful description of the world should be invariant to. We develop these two themes in a domain that forms the basis for widely-used applications, and where ground truth is unusually clear-cut: descriptions of walking routes through real cities. Using a benchmark of 6032 pedestrian corridors in 35 cities, we evaluate the discrete claims made in the descriptions against map data for ten models across families and sizes, and we find initial error rates on landmark claims between 36% and 72%, with consistency across models in which types of errors are the most frequent. Claims about the existence of landmarks are almost always right; models rarely invent a place outright, but they do misplace it. They name real landmarks from the right part of the city but from outside the stretch being walked, and side-of-street claims from the perspective of the walker are essentially random for smaller models. The failure is one of precision rather than recall, and increasing model capability does not eliminate it: the strongest model we evaluate still gets more than a third of its landmark claims wrong. We then study iterative critique-and-revision as a remedy. We formalize this with a three-parameter model: an initial error rate , a per-revision repair probability , and a per-revision corruption probability . The interaction of and produces an equilibrium minimum error rate that revision cannot improve on. Models differ here by degree rather than by kind: rises with capability from about , where revision merely churns errors, to about , lowering the floor from roughly to but never near zero. A city's difficulty is well predicted by two axes that we can measure independently: how much text about the city exists on the Web, and how regular the city’s street network is. For path-dependence, we show that describing the same corridor in reverse leaves average accuracy unchanged but changes the description itself: the two directions name overlapping sets of landmarks with Jaccard similarity only 0.12–0.35, and for landmarks named in both directions the stated side of the street fails to invert between 30% and 52% of the time. A substantial part of this divergence can be seen even in synthetic cities where every relevant fact is supplied in the prompt and no world knowledge is required, which lets us separate a knowledge-asymmetry component from a pure position narration component. We find evidence for this kind of path-dependence even in stylized narrations in other domains, where for example reordering the categories in a household budget changes how much money the model says the family should be saving. We conclude by discussing what these regularities suggest about how such models represent and deploy knowledge of the world, and what they imply for systems that put model-generated descriptions in front of users.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.