acceptodds
Under review as a conference paper at ICLR 2027

When Are Local Linear Models Faithful to Large Language Model Behavior?

Abstract

Local surrogate explanations are used to summarize how black-box models respond to interpretable changes in their inputs. Their fidelity is usually assessed by how well a fitted surrogate predicts the observed responses. Prediction error alone does not distinguish parameter mismatch from model-class limits. We give exact limits for local affine explanations and separate parameter mismatch from the minimum error on the same measured responses. We change one aspect of a request at a time in two opposite ways while keeping the others fixed. For each pair, an affine model can fit the difference between the two responses using its slope, but the average of the two predictions is fixed by the same shared intercept. Therefore, if different pairs have different observed averages, no choice of affine parameters can fit them all exactly. We show that the variance of these averages is the minimum mean-squared error over all affine fits, while their smallest enclosing radius is the minimum error required to fit every measured response. The remaining mean-squared error of a fitted model separates exactly into intercept and slope mismatch. Two response pairs also give a four-response lower bound on the minimum worst-case error. Across 22 language models and eight task-specific neighborhoods, every evaluable setting exceeds the specified affine tolerance under two scorers, including most cases with full-rank directional responses. Across 155 evaluations, the four-response bound reaches a median 96.3% of the exact minimum.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.