acceptodds
Under review as a conference paper at ICLR 2027

Operators Gate Gradients: Routing, Survival, and Reasoning-Shortcut Control in Differentiable Logic

Abstract

A logical constraint can make a network learn less than it would without it: on MNIST digit-sum with the Gödel or Goguen implication, concept accuracy stays near (chance ) through labels, below the supervision alone reaches with . We trace this to the operator that combines sub-predictions, usually chosen by convention. It is credit assignment on two axes: gradient must be *routed* to a parameter, as the computation graph dictates, and *survive*, staying nonzero where training operates, as the objective and grounding dictate. No operator is universally best, but the best is predictable. Fuzzy description logic makes this exact: one implication gates every axiom through closed-form partials, so the chain rule predicts where each operator acts. Across runs, the implication explains – of operator-induced variance once structure is fixed, against – for the t-norm, even under operator-independent scoring; the Łukasiewicz implication is the safe default for pure satisfaction until, as predicted, the objective or grounding changes. Beyond logic, the same axes predict seven of eight aggregation tasks, the one miss being prospective. The paradox is a survival failure. Gödel and Goguen *saturate*, going flat on satisfied constraints, so a wrong-concept solution satisfying them, a *reasoning shortcut*, is a zero-gradient optimum, walled off from supervision by the violated-constraint loss. Swapping in a non-saturated implication lifts the barrier without extra labels, matching supervision by to : the shortcut is a property of the loss, not a shortage of labels.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.