acceptodds
Under review as a conference paper at ICLR 2027

The Signal Survives, the Threshold Slips: Small LLM Agents After Quantization and Fine-Tuning

Abstract

Small language model agents are usually quantized before they are deployed, and fine-tuned so they can call tools. We ran both changes together as a full 2×2 design on six instruction models from 0.5B to 3.8B, and scored what the agents did rather than what they said, using the executable actions in AgentHarm. Across models the interaction term is all over the place, from +0.25 to −0.42, which is easy to read as no effect at all. We show it breaks into three parts once you measure the model's own decision log-odds, meaning the log-odds of opening a tool call rather than refusing. One part is the bounded scale, where an additive shift looks super-additive near the floor and sub-additive near the ceiling. One part is associated with capability loss, which is where the combined pipeline sometimes breaks the agent's ability to emit a well-formed call, and a model that can no longer act gets scored as a safe one, the capability-grounding argument made before us by Vishnubhotla et al. (2026), which we quantify here rather than claim first. What is left is a latent-scale interaction in the decision axis, and on three of the six models it is the largest of the three parts. On Qwen2.5-1.5B and Gemma-2-2B, a one-feature logistic fitted on three conditions predicts the harmful action rate of held-out conditions of the same model without refitting, whether the change came from quantization, an adapter, a learning rate, a seed, a training checkpoint, or a vector added to the residual stream. On the three models where we measured it, the refusal direction barely moves through any of this. It keeps pointing the same way and keeps telling harmful requests from benign ones, while the decision variable shifts about three times further and pushes most harmful requests past the point where the model prefers to act. The safety signal is not being erased. What moves is the threshold.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.