acceptodds
Under review as a conference paper at ICLR 2027

Whose Word Survives Compression?

Abstract

Post-training quantization can preserve answer accuracy while changing the computation that selects between retrieved context and parametric memory. We test causal transfer of source-arbitration interventions under GPTQ INT4 with JuICE, a dual-run attention-head intervention, on Llama-3.1-8B and Gemma-2-9B base. A question-aware matcher audit reveals that apparent transfer can be a same-entity wrap artifact rather than a transferred causal object. On Llama, the original-matcher (un-audited) context-K10 transfer (+15 percentage points (pp), 95% CI [+8, +23]) is one such wrap: the audited number is +1 [+0, +3] (null), with 17 of 17 original INT4 flips naming the same context entity. The locked Llama parametric pair (clean K2) remains detected at n=100 on both precisions under gold-PAR-only counting (−8 [−14, −3] on FP16; −7 [−12, −2] on INT4), confirmed head-specific by a matched random-K2 control. On Gemma, identified context-suppression H− remains detected versus a locked random-K control but is attenuated on INT4 (inclusive −19 [−27, −11] vs. −1 [−5, +2] on FP16; −11 [−18, −4] vs. −1 [−5, +3] on INT4); context-K10 enlarge is a small real effect (+7 [+2, +12] / +5 [+1, +9]) but not INT4-specific, and the original-matcher (un-audited) raw peak L27_H6 is dead as a wrap artifact. Conversely, Gemma's clean parametric pair attenuates to an in-cell transfer null on INT4 despite a detected FP16 effect—the mirror image of Llama's result. A same-sign ∆ is therefore not evidence of transfer without question-aware matching, entity confirmation, and a matched random-head control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.