acceptodds
Under review as a conference paper at ICLR 2027

Finding the Neurons That Ask: Relation-Selective Privacy Control in Fine-Tuned LLMs

Abstract

A language model fine-tuned to extract personally identifiable information answers a request for any relation it was trained on, including relations a deployment must withhold. However, MLP neurons are polysemantic, so imprecise localization damages non-target relations. This work proposes Question-Contrast Average Precision Scoring (QCAPS), which targets the neurons that route a request rather than the knowledge it retrieves. Each neuron is scored by how well its activation separates the target question from contrastive questions over the same context. Each selected neuron is then driven below its off-state by a multiple of its own activation gap, with the multiple calibrated jointly with the selection threshold. Across five relation types on two fine-tuned extraction models, the largest overlap between any two relations' neuron sets is 71 to 250 times lower under QCAPS than under answer-presence identification. The target is suppressed to between 0.6% and 19.2% of baseline, under 10% on eight of the ten model and relation pairs, at lower cost to the remaining relations than parameter-level unlearning on nine of the ten. At a matched neuron budget, substitutions drawn from prior work leave the target at or near baseline. The intervention acts at inference time and is reversed by detaching a hook, without monosemantic decomposition or retraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.