Interpreting Sparse Autoencoder Features in Cytometry Foundation Models through In Silico Perturbations
Abstract
Cytometry foundation models learn sample-level representations for patient-level prediction, yet how they encode immune cell composition remains unclear. Sparse autoencoders (SAEs) expose individual features in these representations, but do not by themselves explain which composition changes those features respond to. We introduce a perturbation-based framework that interprets SAE features through controlled exchanges of cell populations. Resampling a sample’s own cells varies a target population within a specified group while holding cells outside that group fixed. Each annotation specifies the populations exchanged and the direction of the feature’s response, making its interpretation explicit and testable. Across GPCT and MAESTRO, more than 98% of annotations retain their response direction with corrected significance on held-out samples. Annotations of case-enriched features recover 82% and 65% of expert-expected population changes, respectively, across 20 case–control comparisons, while MAESTRO decoder steering reproduces 93.2% of annotated directions. The annotations also support classification from feature descriptions alone. Beyond individual annotations, comparing perturbation responses between cases and controls reveals shared-feature programs linking populations through state-dependent model responses. Together, this framework connects learned features to interpretable composition changes and provides a way to examine how cytometry models represent immune variation across samples and disease states.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.