When and What to Price against Model Extraction
Abstract
As AI systems and web-based services become more capable, an attacker can collect API responses and use them to train a substitute model, reducing the API provider's competitive advantage. A natural defense is to charge more for queries that are useful for extraction, but these charges also affect benign users. In this paper, we study when request-local pricing can provide the desired protection without imposing too much cost on benign users. For any chosen set of queries, we find the rule that makes their answers unprofitable to acquire at the lowest benign burden, using only observable features. With random responses and purchase decisions made before observing them, protection depends on expected charges. Rewriting imposes minimum charges on substitutes, while shared features can pass these charges to benign queries. A linear program finds the cheapest rule for each protected set; choosing among these sets yields the protection–cost frontier. Frontier-Guided Pricing (FGP) uses this frontier to construct candidate rules, then checks their protection and benign burden on independent data. It deploys the qualifying rule with the smallest burden upper bound, or abstains. Across 15 student-capability targets in five extraction settings, all nine deployments meet their prespecified mean targets on held-out runs; no candidate qualifies in the remaining six cases. The frontier thus guides both how to price and whether pricing is worth deploying.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.