acceptodds
Under review as a conference paper at ICLR 2027

On the Geometry of Hierarchy: An Empirical Study in Vision and Language Transformers

Abstract

Does hierarchy in the data become hierarchy in the representation, and if so, in what geometric form? The taxonomies that label vision and language datasets are often hierarchical, but a network trained on them is never told that a beagle is a dog, an animal, or a living thing; so whether a model's representations carry the taxonomic tree, and in what shape, must be tested empirically. We define two measurements: whether a child concept's subspace is contained in its parent's, and whether the parent's feature fires concurrently with the child's. In small transformers trained on synthetic data, tree-structured training data gives rise to nested subspaces, whereas data without hierarchical structure does not. Similarly, across all families of pretrained large vision, language, and vision-language models (VLMs) we tested - examining the image encoders of the latter - hierarchical nesting is present, aligns with the ground-truth semantic tree, and varies little with model size. At the same time, the readouts across models differ depending on the training objective. VLMs (e.g., CLIP) preserve the parent firing on the child at the output, ImageNet training with class labels (e.g., ViT) preserve a portion of it, and self-supervised models (e.g., DINO) preserve almost none of it. Our results indicate that semantic hierarchies in the training data are naturally encoded in nested model subspaces rather than in single directions; however, how much of this hierarchy stays readable differs with the training objective.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.