Does Neighbour Relatedness Matter? Dissimilarity-Priority Packing Contradicts the Similarity Hypothesis in Supervised Fine-Tuning
Abstract
Sequence packing—concatenating multiple fine-tuning examples into one fixed-length training sequence—is standard practice for supervised fine-tuning (SFT) of large language models. A recent line of work argues that pack composition matters and, specifically, that packs of related examples (high embedding similarity) improve downstream quality: Threshold Filtering Packing (TFP) reports gains from packing semantically similar neighbours together. We test this similarity hypothesis directly by proposing and evaluating the opposite policy, dissimilarity-priority packing (DPP), which greedily fills each pack with mutually dissimilar examples via farthest-point selection over example embeddings. Under a matched harness (same model, data, optimizer, step count, and an exactly matched token budget across all policies), we train Qwen3-0.6B and Qwen3-1.7B with four composition policies spanning the full similarity spectrum: DPP (dissimilar), random packing, TFP (threshold-similar), and a nearest-neighbour homogeneous control (HNP). Contrary to the similarity hypothesis, held-out response NLL is a monotone function of realized intra-pack similarity: DPP achieves the lowest NLL (1.5172±0.0014), random is next (1.5189±0.0005), a faithful re-implementation of TFP with its diversity filter intact sits mid-axis at 1.5245±0.0031, and both homogeneous policies (a nearest-neighbour control and an inverted-constraint variant that removes TFP’s diversity mechanism) are worst (1.5395–1.5413, a 0.024-nat gap to DPP). The harmful end of the axis is extreme homogeneity, and TFP’s “maintain sufficient diversity” mechanism is load-bearing: removing it triples the penalty. A block-diagonal attention control shows the effect is entirely mediated by cross-example attention: when leak is masked, all composition policies collapse to statistically indistinguishable quality. Composition choice is thus a real, measurable training signal—but its direction is opposite to the one the similarity hypothesis predicts, and its damage concentrates at the homogeneous extreme. We release our harness and discuss implications for packing-method design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.