IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing
Abstract
Analog in-memory computing (IMC) offers a promising path toward energy-efficient large language model (LLM) inference by executing matrix multiplications (MatMul) directly within memory arrays in the analog domain. Its efficiency, however, comes with an additional source of error: limited-precision analog-to-digital converters (ADCs) quantize the accumulated analog partial sums and introduce output-side error, which is a different problem from conventional activation and weight quantization at the MatMul inputs. Clipping can mitigate both operand and ADC quantization errors by reducing their dynamic ranges, but the optimal clipping factors must jointly balance activation rounding and clipping, weight rounding and clipping, and ADC quantization. Existing clipping methods, designed for digital quantization, do not explicitly optimize this coupled IMC error. Common approaches rely on costly search-based calibration, leading to suboptimal accuracy and long calibration time. We introduce **IMC-CLINIC** (**C**oupled **L**oss-**I**nformed **N**ewton **I**terations for **C**lipping), a clipping calibration framework built around an analytical surrogate for IMC MatMul output error. The loss surrogate models operand quantization, accumulated clipping-induced bias, and ADC quantization jointly, allowing its gradient and approximate curvature to be evaluated directly from only a small calibration set. **IMC-CLINIC** then jointly optimizes activation and weight clipping factors using a safeguarded Newton-type method. We show that **IMC-CLINIC** reduces analog-IMC MatMul output error by effectively balancing operand and ADC quantization errors. Across multiple models and datasets, it improves average zero-shot accuracy by 6.5–11.5 percentage points over the grid search baseline while reducing calibration time by 10.0×–12.1×. Moreover, its analytical surrogate closely tracks empirical IMC output error, while its optimizer is fast and certified within 1% of the global optimum under the loss objective across all projections on two representative models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.