acceptodds
Under review as a conference paper at ICLR 2027

What Does a Recovered Rule Explain? Diagnosing Which Relational Concepts a Neural Model Composes

Abstract

We present a structural explainability method that builds partial, testable explanations of a neural model's decisions and audits where they hold. A symbolic program recovered from a model's decisions alone serves as an initial explanatory hypothesis; the audit then tests its relational constituents against hand-built scenes placed at each constituent's own decision boundary, verified with the task's executable ground truth. Where a constituent fails this test, an explicit alternative program offers a candidate account of the discrepancy, and we evaluate its predictive validity in a held-out case study rather than by how well it matches the errors it was chosen to explain. A linear probe further asks whether the information a constituent needs is at least present in the model's own representation. The result is an explanation profile: a record of which parts of a recovered explanation hold up, which are better accounted for by a named alternative, and which remain unexplained. Applied to Kandinsky Patterns across eight relational tasks and three architectures, including a perception-free model given exact object descriptions, recovered programs agree with the generating rule on standard test examples but fail to predict the model at these targeted boundary cases. For one touching-and-counting task, an alternative that counts touching clusters rather than touching objects raises held-out agreement with the perception-free model from 45.7% to 61.0%; it does not transfer to either image-based model, and matching the model's decisions better than the recovered rule does is evidence about behaviour, not about what computation the model implements internally. Separately, per-object contact remains linearly decodable from the perception-free model's representation at 94.5% accuracy despite the model failing the composed task — a result the probe alone cannot resolve into use or disuse of that information. Rather than a single verdict on whether a recovered rule explains a model, the method produces a structured account of where it holds, where an alternative improves it, and where the evidence available from behaviour and representation runs out.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.