acceptodds
Under review as a conference paper at ICLR 2027

Wrong Rules, Right Predictions: Concept Remapping and Repair in Neuro-Symbolic Models

Abstract

Neuro-symbolic models combine neural perception with symbolic knowledge, but what happens when the symbolic knowledge itself is wrong? A model can compensate for incorrect rules by changing the meaning of its learned concepts while preserving task performance. On MNIST addition, permuted rules retain 97.5% label accuracy even as joint concept accuracy falls to 35.8%. We study what can be certified about such misspecified knowledge before training and whether these guarantees translate into recovery during learning. Given a reference for the intended semantics, we certify whether concept remapping can compensate for an incorrect program and whether the intended program is its unique minimum-edit repair. For full-support addition with concepts, we prove that the nearest competing zero-loss program differs in exactly table entries, yielding a tight unique-repair guarantee for up to corruptions. Yet identifiability does not imply recoverability. On ten previously unseen identifiable twenty-edit tables, a residual encoder recovers the intended program in 48/50 MNIST runs but only 10/50 SVHN runs, while a smaller CNN recovers 46/50 on both. Failures often coincide with the collapse of a single-digit concept. These results separate structural compatibility, repair identifiability, and neural recoverability under misspecified symbolic knowledge.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.