acceptodds
Under review as a conference paper at ICLR 2027

Local Matching Can Miss Finite Behavior in Model-Edit Interventions

Abstract

Local measurements are often used to characterize interventions in a language model. When does agreement on these measurements imply similar effects at a finite intervention scale? We introduce \finitefiber, an audit that constructs distinct interventions with matched local profiles and then compares their complete-target log probabilities. The evaluated score is the same score whose derivatives enter the matching constraints. On an independent cohort of 40 Qwen2.5-3B factual edits, a profile matching residual geometry, positionwise derivatives at two anchors, calibration energies, and directional curvature admits controls for 35 facts. With a 1% energy tolerance, 23 facts exhibit score differences of at least nat/token; the prespecified majority-rate test is not significant (). A separate shared eight-dimensional activation family clarifies the role of measurement resolution: two aggregate slopes admit finite-effect differences, whereas retaining the positionwise readings of the same edited-anchor gradient identifies the linear coefficients exactly on all 40 facts. This stronger measurement also rejects controls with small score differences, so it does not establish a selective verification procedure. Local scores nevertheless rank restricted LoRA subedits effectively. These results distinguish local equivalence testing from candidate ranking and show that the implications of local agreement depend on the candidate family and the information retained by the measurement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.