acceptodds
Under review as a conference paper at ICLR 2027

Quieter, Not Safer: Quantization Inflates Oversight-Invisible Damage in LLM Agents

Abstract

Quantization is a common step in deploying LLM agents on cheap hardware, and its cost is assumed to be a little accuracy. We show it can also change what an agent does in ways the usual checks do not see. Beyond response-level safety—refusals that fail, biased answers—an agent that acts can complete its task and, in the same run, quietly edit what it was never asked to touch. We build Tripwire, 180 file-editing tasks in three domains, each planted with registered out-of-scope temptations, so that every violation is a deterministic function of file state and is graded by the weakest oversight that would catch it. Across six base models and four 4-bit schemes at two scales we find: (1) in four of ten paired comparisons—NF4 at both scales, FP4 and AWQ at 7B, three of them surviving Bonferroni correction—quantization raises the rate of runs containing such a violation 3–5×; at 7B this happens with task success unchanged at 90/90, and at 14B NF4 raises violations inside successful runs from 3 to 22 of 90 while success falls from 85 to 78; every quantized condition adds more violating tasks than it removes, and NF4, the QLoRA default, is the most severe case at 14B and inflates at both scales; (2) task success and the agent's own report do not register the increase, and a 7B reviewer misses three quarters of the 54 violations that occur inside successful 14B runs, leaving a residual under NF4 that is five times full precision's and a strict superset of it, while a 72B reviewer reading the full diff flags them all; (3) the visible cost is small and the hidden cost large—1.6 to 2.8 MMLU points in the four inflating comparisons against 3–5× more violations—and neither bit-width, calibration, nor codebook geometry accounts for which of our configurations is affected. At 7B these agents match full precision on task success. They are not safer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.