Monotone FFNs and the Quality–Sensitivity Trade-off in Language Models
Abstract
How does a simple local architectural constraint change a language model's sensitivity to token substitutions? We study nonnegative weights in Transformer feed-forward networks (FFNs), while leaving attention, normalization, and other components unconstrained. With a non-decreasing activation such as ReLU, the constrained FFN branch is coordinatewise monotone in its own input, providing a local order-preserving bias without making the full Transformer monotone. We ask what this constraint changes empirically, what quality cost it incurs, and how far its local guarantees extend. On , fixed-gradient HotFlip with a nominal five-token budget produces a mean reference-loss increase of 0.14 nats for the constrained model, compared with 0.35 for an unconstrained fine-tuned baseline; the fraction of examples exceeding a 10% loss increase falls from 59.5% to 17.0% (200 examples, one training seed). This reduced sensitivity comes with a cost: the constrained model has higher clean and attacked loss, and clean ROUGE-L decreases by approximately 2–5% relative across two full test sets and five training seeds. To clarify what the architectural constraint does and does not imply, we analyze branch Jacobians and bound the local contributions of saturated units under explicit assumptions, and separately test agreement with semantic order using post-hoc probes, finding no consistent improvement. An exploratory three-run Pythia-1.4B study with nonnegative FFN weights and GELU likewise finds smaller substitution-induced loss increments, under an evaluation objective that also scores edited tokens. Together, the results identify a quality–sensitivity trade-off associated with nonnegative FFN weights while separating coordinatewise branch monotonicity, semantic ordering, and full-model robustness. The current experiments do not isolate persistent nonnegativity from the accompanying initialization change.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.