MindCraft: A Full-Sequence Flow Theory of How Concepts Take Shape in Deep Models
Abstract
Two views describe how a language model represents a concept: as a linear direction or as a curved manifold. Both characterize a representation once it has formed; neither describes how a distinction is carried through depth or predicts which geometry it will take. We introduce MindCraft, which treats depth as time and follows a one-token counterfactual edit through the full-sequence flow operator, the linearized map from the edited embedding to every token at every layer. Each layer stretches and turns the edit, and how the turning is distributed over depth decides the geometry: one sharp turn yields a linear direction, gradual turning yields a manifold. We show that the flow starts isotropic and sharpens at a rate set by the first layer, and that, under a growth gap between answer-relevant and answer-irrelevant directions, each input's flow is driven toward the task read-out subspace. Across four domains and six models from three families, task-central distinctions branch early while temporal and surface details branch late or never. Steering at the branching layers lowers hallucination on HalluLens by 14.5 points and ambiguous-context bias on BBQ to 0.01–0.04.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.