RealAudit: A Training-Realization-Aware Benchmark for LLM Auditing
Abstract
Methods for auditing learned behaviors in language models are typically evaluated on model organisms that differ in both the behavior they exhibit and how that behavior was trained in. This confound makes it hard to tell whether an auditor has found a genuine behavioral tendency or has exploited artifacts of a particular training run. We introduce RealAudit — a benchmark that separates these factors: we implant the same narrow behavior through several training mechanisms and configurations, and characterize each resulting model by its target strength, off-target leakage, and spontaneous self-report. We then evaluate auditors along three dimensions: _how precisely_ they identify the implanted behavior, whether they still find it _across training realizations_, and _how quickly_ they narrow their hypothesis as they interact with the target model. We find that behavioral strength or leakage alone does not determine auditability: models with the same observable tendency can expose different signals to different auditors, and a behavior's auditability can evolve over training on a different trajectory than the behavior itself. Detecting an anomaly is therefore not the same as identifying a behavior, and auditors must be judged on whether they do so specifically, robustly, and efficiently across a behavior's distinct realizations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.