Beyond Fixed Similarity: Instruction-Following Clustering via End-to-End Partition Generation
Abstract
Multi-goal text clustering requires models to move beyond generic semantic similarity and organize text according to diverse user-specified criteria. However, existing benchmarks primarily assume fixed clustering goals or topical criteria, leaving reasoning-intensive, instruction-following clustering underexplored. We introduce ReasonCluster, an LLM-assisted and human-audited benchmark comprising 28 tasks across multiple domains. Its tasks require models to infer latent textual properties and produce instruction-aligned partitions. Using ReasonCluster, we study two routes toward instruction-following clustering. First, we study embedding-based clustering by comparing encoder-based and decoder-based embedders. Decoder-based embedders provide the strongest representations among the evaluated models, suggesting that generation-capable backbones are promising for capturing instruction-dependent textual properties. Yet all embedding-based pipelines still rely on separately configured clustering algorithms, leaving the final partition sensitive to algorithmic hyperparameter choices. We therefore explore generative clustering, where an LLM interprets the instruction, reasons over the corpus, and outputs a complete partition. Our experiments reveal that generative clustering remains challenging for general-purpose LLMs, which substantially underperform reasoning models and are prone to structurally invalid outputs as input complexity grows. We further show that task-specific post-training that optimizes output validity and partition quality can substantially strengthen this capability, enabling a 14B model to surpass o3 by 3.6 points on held-out evaluation split. These results establish ReasonCluster as a test bed for reasoning-intensive clustering and suggest generative clustering as a promising direction for reasoning-intensive end-to-end clustering.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.