acceptodds
Under review as a conference paper at ICLR 2027

Post-Training Adaptation of LLMs to Stochastic Weight Perturbations

Abstract

Analog in-memory computing (AIMC) performs matrix–vector multiplications inside memory arrays that store weights, avoiding costly weight movement during LLM inference. The price it pays is stochastic noise on every weight that is fundamentally different from the deterministic rounding error in digital quantization. Hardware-aware training can mitigate the resulting accuracy loss, but requires a full training pipeline with escalating costs as model sizes grow. Post-training quantization (PTQ) requires only a calibration set, but fails to address stochastic weight noise. We propose a post-training robustification (PTR) method that adapts weights using a noise-aware objective: expected layer-output distortion, decomposed into weight-error and noise-variance terms. Weight-Range Adaptation for Robustness to Perturbations (WARP) combines per-channel input scaling and weight clamping with exactly-solved thresholds for each weight group. WARP+ adds GPTQ-style error redistribution on top. We evaluate 0.6B to 32B models of the Qwen3 family in an AIMC simulator with a hardware-derived programming-noise model. WARP+ improves accuracy across all model sizes and completes in 3 hours on 4 A100 GPUs for the 32B model. On Qwen3-14B, WARP+ recovers 40% of the accuracy loss, more than 2× the recovery achieved by PTQ, and reaches within 3.5% points of noise-free accuracy. On Qwen3-32B, the remaining gap is % points. Using WARP+ for PTQ by applying rounding error as weight noise yields a Qwen3-4B model with 3-bit weights at 9.2% accuracy loss, less than half of OmniQuant (19.3%) and GPTQ (22.4%).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.