Making Transformers Logical
Abstract
Recent work has shown that transformers operating on logical tasks will often fit a simple function that interpolates between observed examples. This means that they often fail to generalize to unseen instances of logical boolean or arithmetic problems. Generalization on the Unseen (GOTU) has therefore become a stress test for understanding the inductive bias of transformers. In this paper, we demonstrate that poor logical extrapolation is not necessarily an inherent property of transformers, and that the desired behavior can be obtained by directly modifying this inductive bias in the network architecture. We show that including a simple attention mask is often sufficient to enable near-perfect extrapolation on boolean tasks and improves sample efficiency on arithmetic tasks. We show that although a similar bias may be encoded by initializing query and key parameters, a mask-based encoding is preserved during training. We additionally explore how a similar mechanism may be applied in vision. These results show the potential for using mask-based priors to encode logical inductive biases in transformers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.