acceptodds
Under review as a conference paper at ICLR 2027

Towards Co-Discovery By Extracting A Symbolic Representation From A Network Layer: Complexity, Algorithms and Experiments

Abstract

Deep learned models (DLMs) have been shown to be useful in a wide variety of situations. However, most XAI methods provide simple instance-level heatmap-style explanations that neither reveal the computation performed by the deep learner nor allow a human to modify the computation. Given that most DLMs are comprised of simple computational units, we explore extracting out a symbolic representation of the computation performed by the model as a formula in disjunctive normal form (DNF). We then show how the DNF can be used not only for explanation but also for simplification based on fairness/complexity/accuracy and even model editing. To achieve a symbolic representation, we restrict the layer's input to binary inputs often found in graphical data and some architectures. We explore three core co-discovery problems: extraction, simplification, and editing, showing complexity results, algorithms, and experimental results for each.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.