On the Effectiveness of “Harmsinks”
Abstract
Removing harmful signals acquired during training poses a challenging problem for AI safety: even when the responsible data can be identified, filtering it out may discard benign learning signals. As a promising alternative, recent parameter- separated training methods gate parameter updates using data labels. The goal is to encourage mechanistic separation at the parameter level, where a subset of weights is updated only by selected training data (Ghosal et al., 2025; Roland et al., 2026). When ablated at inference time, these parameters no longer influence model outputs, thereby “forgetting” the undesirable behavior. However, it remains unclear whether simple data-gated optimization can truly localize complex harmful information, especially in real data. We call the designated subnetwork harmsink and study the harmsink hypothesis that the concept of harmfulness becomes localized in these designated parameters. Specifically, we extend Selective Gradient Masking (SGTM) (Shilov et al., 2025) to FineWeb data annotated for harmfulness. Prior work evaluates such localization using the loss gap between forget (harmful) and retain (benign) datasets, but this behavioral metric does not establish mechanistic separation. Across four small model scales, layer-wise linear probes show that the harmful–benign distinction remains decodable after harmsink ablation, indicating that mechanistic separation often fails to emerge despite observed loss gaps. We additionally test the proposed parameter auto-routing behavior that motivates this approach and find little evidence for it. Our findings suggest that forget–retain loss differences can overstate parameter separation and motivate further work toward more reliable harmsinks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.