Scale-Calibrated Preconditioning for Layer-Wise Clipping under Differential Privacy
Abstract
Recent advances have improved the efficiency and utility of differentially private training for large vision and language models. Per-layer gradient clipping is an important approach to reducing memory overhead. However, clipping each layer independently with its own threshold can produce different scaling factors across layers, distorting the full-gradient direction and degrading utility. We propose SCP-Clip, which performs per-layer clipping in a scale-calibrated preconditioned space to jointly mitigate directional distortion caused by per-layer clipping. The method incurs no additional privacy cost while retaining the efficiency advantages of per-layer clipping. Theoretically, we establish a utility bound on the private gradient estimation error under stable feedback. Experiments on pretrained language and vision models show SCP-Clip outperforms flat clipping in most settings, with accuracy gains of up to 4.28% over adaptive per-layer clipping.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.