How Extractable Are Edge LLMs?
Abstract
Edge large language models (Edge LLMs) are commonly quantized for deployment on devices with limited memory and compute. At the same time, their text interfaces allow an attacker to collect responses and train a student model, creating a risk of black-box behavioral extraction. For such attacks, collecting more target responses provides additional supervision; however, slow on-device inference makes large-scale data collection costly. We therefore investigate whether an attacker can extract instruction-following behavior more effectively under a limited target-query budget. To assess this risk, we introduce Coordinated Linguistic Instruction Querying (CLIQ), a structured probing framework designed to reduce repeated probing of similar behaviors by coordinating queries across semantic regions. Specifically, CLIQ partitions a public instruction pool and generates region-conditioned queries before observing any target response. We compare it with Original Queries (OQ), which are randomly sampled from the same pool and used without modification. Under a nominal 500-query budget, CLIQ improves a 4B student's BERT-F1 from 77.92 with OQ to 82.44 on Alpaca evaluation references. Furthermore, experiments with two phone-deployed Gemma teachers show improved student performance over OQ while target-side generation speeds remain comparable. Taken together, these results demonstrate that functional instruction-following behavior remains transferable from quantized edge models under the evaluated query budgets: the cost of collecting target responses does not, by itself, establish resistance to behavioral extraction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.