Mitigating Entropic SNR Collapse in Cross-Lingual Differential Privacy
Abstract
Conventional sequence-level gradient clipping and length normalization in Differentially Private (DP) fine-tuning implicitly assume a Latin-centric, homogeneous subword distribution. However, they adversely compound under multilingual distributions because unequal vocabulary allocation inflates sequence lengths for non-Latin and morphologically rich scripts by up to . Under a stylized model of semantic signal conservation, sequence-averaged gradient normalization attenuates signal power quadratically, inducing an upper bound on vector Signal-to-Noise Ratio (SNR) against isotropic Gaussian DP noise, a degradation phenomenon we formalize as Entropic SNR Collapse. Existing solutions face severe trade-offs: sequence-level adaptive clipping fails to regulate heavy-tailed intra-sequence token spikes; per-token clipping with active length normalization collapses parameter update norms by , severely impairing generation fluency; and tokenizer retraining breaks pretrained model interfaces. To mitigate this issue, we propose ToPA (Token-Aware Privacy Alignment), a framework that decouples token sensitivity bounding from sequence aggregation. ToPA enforces localized gradient bounds via unified Token-Wise Ghost Clipping across Transformer projections without materializing per-token Jacobians, accumulates updates via a fixed-anchor aggregator to reflect token-uniform empirical risk without leaking metadata, and privately releases clipping thresholds via the Exponential Mechanism. Extensive downstream benchmarks across four model scales (1B to 9B parameters) on translation (FLORES-200), cross-lingual inference (XNLI), and question answering (TyDi QA) demonstrate that ToPA stabilizes token gradients, prevents optimization collapse on fragmented languages and substantially closes the gap to non-private baselines under strict DP budgets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.