CoClusterPFN: Amortized Posterior Similarity for Tabular Bayesian Clustering
Abstract
Most clustering methods return one grouping without revealing which assignments are uncertain. Bayesian methods represent this uncertainty but typically require expensive inference. We introduce CoClusterPFN, a Prior-data Fitted Network (PFN) for Bayesian clustering of tabular data. It is trained on synthetic clustering tasks to predict pairwise posterior co-clustering probabilities in a single forward pass. The resulting posterior similarity matrix (PSM) avoids the label-switching ambiguity of cluster assignments. When the number of clusters is unspecified, the prediction marginalizes over cluster-count uncertainty. The same PSM provides an interpretable summary and can be decoded into hard partitions using decision rules targeting different metrics or selecting different numbers of clusters. Under a Gaussian-mixture prior, CoClusterPFN closely matches PSMs estimated by Markov chain Monte Carlo (MCMC) while being orders of magnitude faster. On 60 CLUBench datasets, CoClusterPFN achieves the best average rank for adjusted Rand index (ARI), normalized mutual information (NMI), and Binder loss, while ranking second for variation of information (VI), whether the cluster count is supplied or unknown. Beyond benchmark performance, an Iris case study demonstrates how CoClusterPFN can help users choose among alternative clusterings. Together, these results support CoClusterPFN as an effective approach to uncertainty-aware clustering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.