When Are Local Linear Models Faithful to Large Language Model Behavior?
Abstract
Local surrogate explanations are used to summarize how black-box models respond to interpretable changes in their inputs. Their fidelity is usually assessed by how well a fitted surrogate predicts the observed responses. Prediction error alone does not distinguish parameter mismatch from model-class limits. We give exact limits for local affine explanations and separate parameter mismatch from the minimum error on the same measured responses. We change one aspect of a request at a time in two opposite ways while keeping the others fixed. For each pair, an affine model can fit the difference between the two responses using its slope, but the average of the two predictions is fixed by the same shared intercept. Therefore, if different pairs have different observed averages, no choice of affine parameters can fit them all exactly. We show that the variance of these averages is the minimum mean-squared error over all affine fits, while their smallest enclosing radius is the minimum error required to fit every measured response. The remaining mean-squared error of a fitted model separates exactly into intercept and slope mismatch. Two response pairs also give a four-response lower bound on the minimum worst-case error. Across 22 language models and eight task-specific neighborhoods, every evaluable setting exceeds the specified affine tolerance under two scorers, including most cases with full-rank directional responses. Across 155 evaluations, the four-response bound reaches a median 96.3% of the exact minimum.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.