Beyond Permutation Invariance: Idempotence in Set Data Representation
Abstract
Permutation-invariant architectures are routinely described as learning on sets, but permutation invariance removes order without removing multiplicity. This distinction is often hidden because datasets typically provide only one finite registration of each underlying datum. If repeated registration of an already-present element does not change that datum, its representation must also be consistent across such alternative registrations. We formalize this requirement as idempotent set consistency and prove that a variable-cardinality family descends to a function on finite subsets exactly when it is permutation invariant and compatible with duplicate insertion and deletion. This yields registration fibers and quantitative fiber defects that distinguish set semantics from multiset and empirical-measure semantics. More generally, we use directed systems of registrations as a semantic checklist: the chosen cross-cardinality maps specify which changes of registration are intended to preserve the datum, and their compatibility reveals which invariances a representation should satisfy. For set semantics, this perspective leads naturally to Hausdorff geometry and near-idempotence under arbitrarily close observations. We study these ideas through pooling and Set Transformer audits, geometry-controlled interventions on pretrained ModelNet40 classifiers, and duplicate and near-duplicate corruptions in human trajectory prediction. Rather than treating idempotence as a universal desideratum, our framework uses registration consistency to distinguish competing data semantics and tests whether a model respects the one intended by the application.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.