Reading between the Clauses: Data Reconstruction Attack on Federated Training of Tsetlin Machines
Abstract
Tsetlin Machines (TMs) classify binary inputs through learned logical combinations of features. They have recently gained attention because they train without GPUs, run efficiently on edge devices, and are explainable by design: each learned clause is a small pattern that points directly at the input pixels behind a decision. Weighted Convolutional TMs (wCTMs) have even been shown to outperform neural networks on federated image classification under strong data heterogeneity. Federated learning is meant to protect privacy by keeping each client's data on their own device and sharing only the trained model with the server. We show that for wCTMs this protection fails: the same explainability that makes a clause readable also makes it invertible. Under heterogeneous data, many clients hold classes with so few images that their clauses memorize the individual samples. We present the first data reconstruction attack on federated wCTM training. From a single uploaded model, an honest-but-curious server estimates how many images a class was trained on and reassembles the clauses' learned patches to recreate the original images. No auxiliary data and no protocol manipulation are required. Across five benchmarks, the attack reconstructs hundreds to thousands of images, up to 95% of the images each heterogeneous setting leaves exposed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.