Decodability Is Not Transferability: Uncovering Limitations in Graph Attribute Auditing via Orthogonal Alignment and Invertible Rotation Controls
Abstract
Attribute inference auditing evaluates whether sensitive attributes beyond the intended task can be decoded from model outputs or internal representations - an important privacy concern for large language models that process graphs. Existing audits use output confidence or train attribute classifiers on shadow-model representations, assessing attribute decodability through transfer to target representations. However, whether cross-model transferability reliably reflects target-side attribute decodability remains unclear. We systematically investigate this relationship through orthogonal alignment and invertible rotation techniques. Orthogonal alignment preserves attribute information while adjusting shadow representations’ coordinate directions to better match target representations. Under identical preprocessing, alignment improves transfer performance by up to approximately 0.30 macro AUROC, suggesting that poor transfer may partly arise from coordinate mismatch rather than difficulty in decoding target attributes. To further verify this, we use invertible rotation to change only the coordinate directions of target representations without removing information, reducing the shadow-trained classifier’s transfer performance by approximately 0.16 macro AUROC while leaving the performance of a linear classifier retrained on target representations nearly unchanged. Overall, attribute decodability and cross-model transferability are not equivalent: low transfer scores do not directly imply that target attributes are difficult to decode. These findings reveal the limitations of assessing attribute exposure solely through transfer performance. Transfer-based audits should therefore interpret results in light of the classifier’s training source and the characteristics of the representation spaces.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.