acceptodds
Under review as a conference paper at ICLR 2027

Making Transformers Logical

Abstract

Recent work has shown that transformers operating on logical tasks will often fit a simple function that interpolates between observed examples. This means that they often fail to generalize to unseen instances of logical boolean or arithmetic problems. Generalization on the Unseen (GOTU) has therefore become a stress test for understanding the inductive bias of transformers. In this paper, we demonstrate that poor logical extrapolation is not necessarily an inherent property of transformers, and that the desired behavior can be obtained by directly modifying this inductive bias in the network architecture. We show that including a simple attention mask is often sufficient to enable near-perfect extrapolation on boolean tasks and improves sample efficiency on arithmetic tasks. We show that although a similar bias may be encoded by initializing query and key parameters, a mask-based encoding is preserved during training. We additionally explore how a similar mechanism may be applied in vision. These results show the potential for using mask-based priors to encode logical inductive biases in transformers.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.