Language Models That Walk the Talk: A Framework for Formal Fairness Certificates
Abstract
As large language models become integral to high-stakes applications, ensuring their robustness and fairness is crucial. Despite their success, large language models remain vulnerable to adversarial attacks, where small perturbations can alter model predictions. In language models, perturbations such as synonym substitutions also pose risks in fairness- and safety-critical areas. Formal verification is a viable path to certify fairness; however, its application to large language models remains limited. This work presents a holistic verification framework to certify the fairness of language models by formalizing the verification problem within the embedding space, with a focus on ensuring gender fairness and consistent outputs across different gender-related terms. Furthermore, we extend this methodology to toxicity detection to offer formal guarantees that adversarially manipulated toxic inputs are consistently detected and appropriately censored, thereby ensuring the reliability of moderation systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.