Safe in Text, Unsafe in Action: Distinguishing Content and Physical Danger in LLM Hidden States
Abstract
Deploying large language models (LLMs) in embodied agents extends safety concerns from generated content to physical actions. Although content safety has been extensively studied, it does not guarantee physical safety, as seemingly benign instructions can become hazardous when executed in the physical world (e.g., “microwave a metal fork”). We identify content danger (CD) and physical danger (PD) as distinct safety dimensions and investigate whether LLMs internally distinguish them. Across six LLMs spanning 1.7B32B parameters and three model families, we find statistically separable CD- and PD-associated directions in hidden-state representations (, the 999-permutation floor). This separation exposes a gap between internal safety representations and verbal safety judgments. Across four benchmarks, a linear readout of hidden states outperforms the model’s own verbal judgment by 0.030.30 AUC, with the largest gap occurring for physically grounded danger. A rank-16 path from the probed layer into the final decoder block recovers 2475% of this gap without changing any base weight, while the same construction at the embedding layer does not, meaning the signal is present but under-used. A probe transferred from a true/false benchmark reaches only 0.600.66, so it is not generic truthfulness. Building on this finding, we introduce PRISM (Probing Representations for Integrated Safety Monitoring), a lightweight detector that identifies physical risks using a linear probe of a single hidden layer without modifying the base model. Across SafeAgentBench, SafeText, and EARBench, PRISM reduces false positives by 5075% compared with verbal safety judges. These results show that LLMs internally encode physical danger as a distinct safety dimension, but fail to fully route this information into their own safety judgments, enabling hidden-state representations to serve as an effective signal for physical-risk detection in embodied agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.