Uncovering Hierarchical Scene Structure from Human Perceptual Grouping
Abstract
Understanding the hierarchical organization of natural scenes remains a persistent challenge in computer vision. In this work, we present a data-driven framework for recovering multi-scale scene hierarchies directly from human visual grouping judgments without requiring semantic labels, demonstrating that universal grouping principles can be learned independently of semantic category knowledge. We introduce Triple-Seg, a benchmark dataset of 210 natural images paired with segmentations and dense human triplet comparison judgments of the image regions. Using a neural embedding model trained on these judgments, segment-level visual features are mapped into a embedding space that enables the reconstruction of structural scene dendrograms via agglomerative clustering. By relying on perceptual rather than semantic cues, the model generalizes to unseen categories from only a few hundred training images. Code and datasets will be made publicly available upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.