Responsiveness Verification: Will Predictions Change? How Much? How Often?
Abstract
Machine learning models are often used in applications where their inputs change due to routine interactions, strategic manipulation, or noise. In such settings, models can undermine safety as these changes can lead them to predict over regions of the input space they have not seen. We propose to address these challenges by measuring *responsiveness*—the probability that a model output attains a target prediction under feasible interactions defined by the interaction model. We develop algorithms to estimate responsiveness for any machine learning model, and a framework to specify broad classes of interaction models. We pair these algorithms with statistical guarantees that support practical validation. We demonstrate how our tools can promote safety and reliability across domains by detecting preclusion in recidivism prediction, estimating the cost of gaming in content moderation, and testing the robustness of benchmarks for LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.