The Frontier Is Never Understood
Abstract
Whether a model can be completely understood by a model no more capable than itself has been asked informally for years and settled neither way, because neither term has been tied to what a system may actually spend. We fix both to a declared deployment budget: a system is frontier at that budget if nothing within the budget strictly dominates it, and an interpreter understands a system if it retains that system's capabilities and answers a fixed specification of mechanism questions about it. The question then has a two-case answer. Either a frontier system already answers its own mechanism questions, or nothing within its budget understands it, and whatever does strictly dominates it and takes its place; the most capable system in a cohort is therefore unexplainable from inside that cohort rather than merely unexplained. What the escape costs depends entirely on whether the target is fixed when the interpreter is built or supplied with each request. In the first case the optimal interpreter is a lookup table, at a marginal cost of exactly m stored parameters for an m-answer specification, a tight bound on a construction that explains nothing. In the second the parameter cost vanishes and the whole cost falls on the inference deadline, where we bound it and measure it on a transformer across six standard mechanism queries: the factor ranges from one to 2^(LH), and the number of times understanding can be iterated from about eight to fewer than one. The obstruction reaches only overseers that retain their target's capabilities, leaving weaker-evaluator schemes untouched; where it does reach, conditioning deployment on understanding fixes a permanent gap between the most capable system built and the most capable one deployed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.