acceptodds
Under review as a conference paper at ICLR 2027

RealAudit: A Training-Realization-Aware Benchmark for LLM Auditing

Abstract

Methods for auditing learned behaviors in language models are typically evaluated on model organisms that differ in both the behavior they exhibit and how that behavior was trained in. This confound makes it hard to tell whether an auditor has found a genuine behavioral tendency or has exploited artifacts of a particular training run. We introduce RealAudit — a benchmark that separates these factors: we implant the same narrow behavior through several training mechanisms and configurations, and characterize each resulting model by its target strength, off-target leakage, and spontaneous self-report. We then evaluate auditors along three dimensions: _how precisely_ they identify the implanted behavior, whether they still find it _across training realizations_, and _how quickly_ they narrow their hypothesis as they interact with the target model. We find that behavioral strength or leakage alone does not determine auditability: models with the same observable tendency can expose different signals to different auditors, and a behavior's auditability can evolve over training on a different trajectory than the behavior itself. Detecting an anomaly is therefore not the same as identifying a behavior, and auditors must be judged on whether they do so specifically, robustly, and efficiently across a behavior's distinct realizations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.