acceptodds
Under review as a conference paper at ICLR 2027

CrystalSAE: Interpreting and Steering Crystal Diffusion Models through Sparse Autoencoders

Abstract

Crystal diffusion models perform well in crystal structure prediction, but their internal representations remain poorly understood. To improve model interpretability and inform controllable generation, we propose CrystalSAE, a sparse autoencoder (SAE) framework for identifying and probing crystallographic features in crystal diffusion models. Specifically, we apply SAEs to the layer-wise residual updates of DiffCSP, a typical crystal diffusion model, and relate the resulting feature activations to crystallographic concepts. We then intervene along selected feature directions to examine their influence on generation. Extensive experiments on MP-20 dataset confirm that sparse features selected on training data retain statistically significant associations with 17 crystallographic concepts on held-out data. Crystal visualizations further illustrate associations between selected features and 68 space groups. Moreover, our intervention experiments show that stronger perturbations of selected features produce larger average changes in the lattice parameters and density of generated crystals. These findings reveal crystallographically aligned features within DiffCSP and suggest a path toward more interpretable and controllable crystal generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.