acceptodds
Under review as a conference paper at ICLR 2027

Using Rashomon Sets to Reduce Identity Leakage in Data Collected from People

Abstract

Machine Learning models increasingly recognize gestures and activities from behavioral data. The same measurements can also reveal who produced them. We use Rashomon sets and Model Class Reliance to construct models that retain the features most essential to a prediction task while simultaneously reducing reliance on identity-rich features. The models in a Rashomon set (i.e., all near-optimal models) often rely on different combinations of features. Treating identity as a second prediction task, we use Model Class Reliance to find features on which every near-optimal identity model depends, then trace effects of removing those features. We evaluate this approach on 10 public datasets from six sources, covering touch gestures, activity sensing and hospital records. As an example, in 14 of 18 ShearSense leave-one-subject-out folds, every near-optimal gesture tree uses at least one identity-reliant feature, so lower identifiability cannot be obtained simply by choosing another near-optimal gesture model. We therefore remove identity-reliant features one at a time, trace the full accuracy-identifiability curve, and cut it where further removal would cost more target task accuracy than identity accuracy. On the 6 datasets where the cut lowers identity accuracy by at least 5 percentage points, identity accuracy falls by 8.7 to 58.0 points and target task accuracy by 0.4 to 14.1 points. A pre-trained tabular deep model shows a similar target task cost. The tradeoff gives designers explicit control over target task accuracy and empirical identifiability, while identity retained in raw sensor data remains a system-design consideration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.