acceptodds
Under review as a conference paper at ICLR 2027

CONTRACT-ALIGNED QUANTIZATION FOR FSS- BASED SECURE TRANSFORMER INFERENCE

Abstract

Function secret sharing (FSS) makes online two-party inference cheap, yet existing quantized Transformers for private inference are trained with at mostbitwidth-level awareness of the protocol—an idealized forward pass that doesnot match what the deployed protocol gates actually compute: FSS gates impose share-wise rounding with correlated truncation errors, public lookup-table encodings, and gate-specific input domains. We show that this mismatch systematically degrades accuracy. We therefore propose contract-aligned quantization: every nonlinear operator first publishes an explicit operator contract — ring, fixed-point scale, slot domain, LUT payload encoding, and truncation path — and the pruned, quantized model is then retrained with the exact integer semantics of the deployed gates in its forward pass. On SST-2, we retain 98.5% of the teacher's accuracy, Compared item by item with the values reported by SIGMA, Softmax achieves approximately 77% lower latency and approximately 51% lower communication, while GELU achieves approximately 55% lower latency and approximately 96% lower communication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.