GIGA: Grouped Sparse Updates for Large-Batch Recommendation Model Training
Abstract
Scaling the worker count at a fixed local batch size increases throughput but also coalesces sparse updates, which in turn often degrades downstream recommendation quality. In click-through rate (CTR) models, this effect is particularly frequency-sensitive: frequent IDs lose more update opportunities over a fixed sample budget than rare IDs, so a single sparse learning-rate multiplier cannot simultaneously correct head and tail IDs. To this end, we propose **Group-wise ID Gradient Aggregation (GIGA)**, which restores fine-grained control over update coalescing. Specifically, GIGA averages each active ID's gradients within worker groups, sums the group means for the update numerator, and separately accumulates their elementwise squares. Group size interpolates between active-worker averaging and retaining every active-worker contribution, while leaving single-worker IDs unchanged. GIGA requires neither historical frequency statistics, additional collectives, nor persistent per-ID state. We show that, for fixed local batch and group size, the expected number of retained group terms over a given sample budget is independent of worker count, and relate the fused update to sequential optimization. In a controlled 96-worker comparison, GIGA improves AUC and request-level GAUC by and percentage points, respectively. Combined with dense-parameter scaling, GIGA supports a 16-fold larger global batch of 307,200 examples on 512 workers and delivers sample throughput while matching or improving both metrics relative to the 32-worker baseline. Gains extend to a separate search CTR model under warm-start and from-scratch training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.