Anatomy of a Jump
Abstract
Can a language model infer an explanatory concept that the observations do not name, rather than fit a familiar law to the state it is shown? We test this in 48 procedurally generated physical worlds. At the restructuring level, the best tested local patch stays below on the certification samples while the true law reaches 1.000; an ideal observer selects the true structural class after one rendered trajectory when that class is supplied among three candidates. Twelve model configurations face a capability check, and two run the full ladder, producing 1,872 answers, although only the frontier arm passes the check. The paper presents three results. First, all 893 parseable stated laws use only position, velocity and time, but the answer format required those variables. Without that restriction, four of seven models still produce no uncued law reading additional state; the other three produce one each, two by split votes. A cue elicits hidden state in 60 of 504 laws, none the world's governing law. Second, in 1,223 of 1,336 scored answers, the model's numerical prediction beats its own stated law, including 398 of 445 at retrieval, where the correct law fits the format. Third, a four-lag linear autoregression fitted to the same fifty prompt rows beats the frontier model on 124 of 153 powered restructuring trials, with roughly eighteen times smaller median error. What this control licenses is a narrower rule: predictive accuracy can support a world-model inference only where ontology-free predictors fail on the same regime. One of three families qualifies over the swept thresholds; its advantage has the same sign on fresh worlds, but both intervals include zero and the qualifying partition changes. Mechanistic probes at 1.7B yield bounded null results.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.