acceptodds
Under review as a conference paper at ICLR 2027

COQUAD: CO-OPTIMIZED QUADRATIC ACTIVATIONS FOR EFFICIENT PRIVATE TRANSFORMER INFERENCE

Abstract

Ensuring private inference for Transformer architectures remains challenging, with nonlinear activations serving as a key bottleneck. Quadratic replacement offers inexpensive secure evaluation, but recovering model accuracy requires adaptation. We show that the role of quadratic coefficients depends on the FFN architecture: they can be absorbed into adjacent affine maps in biased standard FFNs, whereas nondegenerate bias-free gated channels retain a coefficient ratio invariant under projection reparameterization. MPCFormer's fixed polynomial prescribes this ratio; CoQuad allows it to adapt by jointly learning layer-wise quadratic coefficients and network weights. DirectFit supplies data-dependent initialization, and Co-Optimize performs teacher-guided joint adaptation while retaining quadratic secure evaluation. On ImageNet-1K, full replacement in ViT-B/16 achieves 79.83% Top-1 accuracy (1.62 percentage points below the 81.45% teacher), exceeding a ViT adaptation of the full MPCFormer method by 1.42 percentage points. Llama experiments extend the evaluation to bias-free SwiGLU: with the same 16 layers replaced and the same 1,500 training steps, CoQuad reaches 72.6% ARC-Easy accuracy, compared with 36.5% for MPCFormer's fixed polynomial. Compared with SHAFT, quadratic GeLU evaluation is estimated to be 47–59× faster and reduces communication by 62×; with the other operators unchanged, measured end-to-end speedups over SHAFT range from 1.227× to 1.530×. These results demonstrate the practical value of joint coefficient and network adaptation across standard and gated Transformer FFNs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.