Shape-Selective Correlation Conversion for Public-Shape Private Transformer Inference
Abstract
Private inference systems typically prepare correlations for a fixed tensor graph, whereas token pruning and elastic networks may reveal a public tensor shape only after preprocessing has been distributed. Preparing every supported shape duplicates correlations, while generating and delivering preprocessing after disclosure delays online execution. We introduce shape-selective correlation conversion (SSCC), which locally specializes distributed preprocessing while preserving the fixed-shape online protocol. For retained bilinear correlations, we identify the missing contraction term and characterize the exact minimum additional ring elements required by zero-interaction share-linear decoders over finite chain rings. We instantiate SSCC for nested attention capacities and elastic SwiGLU widths and compose both axes in one Transformer block without a capacity-by-width preprocessing family. For public shape selection and one-use preprocessing, conditional completion lifts fixed-shape security under branch-isolation and joint-pseudorandomness assumptions. Across nine capacity-width pairs, selected and separately generated fixed-shape outputs agree exactly and incur identical online communication. Factorized preprocessing reduces per-party storage by 50.39% relative to a capacity-deduplicated fixed-shape baseline, and a synthetic 12-block chain checks cross-block output and communication consistency. A BERT-base case study on the Microsoft Research Paraphrase Corpus matches the unpruned accuracy of 86.52%. Across 20 co-located repetitions per method and width excluding delivery, the observed median interval from public width disclosure to output is 16.3% to 20.7% lower than for generating fixed-shape preprocessing after disclosure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.