J-FDCS: Jacobian Fault Diagnosis Concept Spaces for Interpreting Large Language Models
Abstract
When a language model diagnoses a machine from vibration signals, it remains unclear whether it represents the underlying fault or merely learns to produce the expected answer. Diagnostic accuracy alone cannot resolve this distinction. We introduce Jacobian Fault Diagnosis Concept Spaces (J-FDCS), a framework for examining fault-related representations in signal-grounded language models. Models receive vibration signals with equipment and task context and learn to answer with abstract codes. We analyze their hidden states using natural fault expressions. J-FDCS constructs expression-associated directions through the average Jacobian Lens or a full-sequence gradient formulation. It then removes their linear overlap with specified task-control and answer-code directions. The resulting low-dimensional space supports complementary tests of readability, semantic correspondence, signal dependence, cross-state stability, and local output effects. We systematically evaluate Qwen, Llama, and Mistral on five public rotating-machinery benchmarks and an in-house BW250 drilling-pump vibration dataset. The results distinguish strong linear class decodability from the variable readout of natural fault expressions. Fixed-space perturbations reveal input-responsive readouts, while cross-state transfer and retrospective BW250 interventions show variable semantic retention and local changes in diagnostic answer scores. Beyond predictive accuracy, J-FDCS provides a way to examine what fault-related information a model represents, which signals it depends on, and how changing it affects diagnosis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.