acceptodds
Under review as a conference paper at ICLR 2027

When Are Safety Anchors Enough? Availability and Certified Preservation

Abstract

Safety anchors constrain fine-tuning, but a useful guarantee must connect available example gradients to the updates a provider actually permits. We study preservation under a single global parameter update chosen from an enforced region. A spectral nonlinear risk gate turns available-span geometry into explicit pool-expansion, constraint-budget, and abstention decisions. For noncentered Gaussian gradients, an exact scalar representation connects systematic drift to the same spectral feasible families. A complementary pointwise envelope supports conditional flip-rate calibration through fixed-sequence testing. We characterize why this envelope can be vacuous even when every global update is safe, and why average-risk control can permit complete failure on a rare category. Experiments test unplanted 128-dimensional pools, conflicting nonlinear tasks, disjoint handwritten-digit splits, and global curvature bounds through three hidden layers. The results distinguish larger retained-direction displacements from improved task progress: anchors can enable a prescribed protected direction yet greatly limit progress on a conflicting task. On real classifiers, a 2% conditional flip ceiling gives modest task gains and shows clear calibration and curvature limits. The contribution is an anchor-specific availability and certification analysis, with explicit successful and unsuccessful operating regimes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.