acceptodds
Under review as a conference paper at ICLR 2027

Group-Balanced Supervised Fine-Tuning For Knowledge-Intensive Domains

Abstract

Supervised fine-tuning (SFT) weights every supervised token equally, so on knowledge-dense corpora the loss budget follows token count rather than importance: on a legal QA benchmark, citation strings absorb 81% of the supervised loss weight budget while the answers being taught absorb 5%. We present Group-Balanced SFT (GB-SFT), which equalizes the loss budget across meaningful groups of tokens, as a two-parameter objective family that contains standard SFT, class-level balancing, and span-level balancing as corners, whose geometry we characterize exactly. Across four benchmarks in law and medicine on two model families, we show that: (i) GB-SFT injects new knowledge better than uniform SFT, pointwise token weighting, memory editing, and matched-budget data mixing; (ii) it matches the injection of masking out everything but the knowledge payload while retaining the citation recall that masking destroys (0.98 vs. 0.38); and (iii) role-aware grouping carries the gain: label-free partitions recover at most a third of it, while an LLM tagger recovers nearly all, but only when the groups are lexically distinctive. We release two new contamination-controlled knowledge-injection benchmarks with validated ground truth.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.