acceptodds
Under review as a conference paper at ICLR 2027

Benchmarking Interpretability Localizers For Adaptation Placement

Abstract

Parameter-efficient fine-tuning makes it cheap to adapt a language model to a shifted distribution, but not to decide which layers should receive the update. Mechanistic-interpretability localizers—sparse autoencoders, attribution patching, gradient and Fisher saliency, drift-native gradient directions, linear probes—promise "where," and that promise is a falsifiable prediction: if a localizer identifies where the shifted computation lives, adapting those layers should beat adapting anywhere else at the same budget. We build the ground truth to test it: dense grids of matched-budget LoRA adapters at every candidate layer window, across four domains and three model families (Gemma-2 9B; Qwen-2.5 7B and Llama-3.1 8B), each grid measuring which placements actually recover the task. Scored against the grids under a frozen protocol, seven localizers—including AdaLoRA, a method built to allocate adaptation budget—yield one verdict: no localizer reliably predicts where to adapt. Where placement is certifiably flat, every localizer ties a random matched-budget window; where placement demonstrably matters, even the localizers that name the peak confer no certified advantage over a random window; a depth band frozen on two model families breaks on the third; and no windowed placement, localizer-chosen or optimal, recovers whole-network performance. The negative is bounded, not merely unresolved: no tie conceals a localizer advantage above +5.7pp (95% confidence). We release PlacementBench—the grids, the frozen scoring-and-admissibility protocol, and effect bounds on every tie—as an instrument rather than an indictment: a future localizer overturns this result by clearing its two bars, identification and certified advantage. A first use: a cross-cell readout calibrated on the graded grids identifies a top-2 window in 8/9 cells and clears the certification bar on one, where every per-cell rule fails.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.