acceptodds
Under review as a conference paper at ICLR 2027

Faithful Rules Are Not Enough: Boolean Surrogates for Deep RL and an Audit of Their Concepts

Abstract

Rule surrogates are among the most readable explanations of deep reinforcement learning (DRL) policies, and their quality is almost always judged by a single number: how often the rules reproduce the policy's actions. We ask whether that number says anything about the concepts the rules are written in. We first build a strong surrogate for vision-based policies. LUCID jointly trains a sparse autoencoder on the frozen policy's features, a learnable binarization bottleneck and a product t-norm logic layer, and extracts Disjunctive Normal Form (DNF) rules; once the selectors polarize, the extracted rules are invariant over a wide range of extraction thresholds, and a state-level certificate bounds the disagreement between the rules and the continuous surrogate. Across five vision-based tasks, from MiniGrid to Pong and Boxing, the rules agree with the PPO teacher on 96.8–100% of states and need about 2× fewer feature tests than a decision tree of matched fidelity on gridworlds and 28–46× fewer on Atari; controls show that joint training is what keeps them faithful after extraction. We then audit the concepts along two axes: alignment with simulator ground-truth factors, and causal effect under intervention on the frozen teacher. Under our criteria the two never hold together in any of the five environments, and on Dynamic-Obstacles rules with near-perfect fidelity rest on concepts whose manipulation moves the teacher away from the predicted action. VLM-generated concept labels, tested blind, predict their concepts above chance but well below certainty. Behavioural fidelity, factor alignment and causal efficacy are distinct properties, and rule-extraction work should report all three.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.