acceptodds
Under review as a conference paper at ICLR 2027

POISONING THE FIT SET: END-TO-END ATTACKS ON OPQ-BASED SEMANTIC-ID RECOMMENDATION

Abstract

Semantic-ID generative recommenders often rely on an OPQ tokenizer fitted on a subset of catalog items and then used to assign discrete identities to the full catalog. We identify this fit set as an overlooked trust boundary and ask whether injected catalog items can alter the Semantic IDs of legitimate, untouched items and affect downstream recommendation. Clean refitting is exactly reproducible, with Adjusted Rand Index (ARI) 1.000, allowing structural changes to be attributed to fit-set modification rather than tokenizer randomness. Across three OPQ-based recommendation pipelines and three Amazon catalogs, fit-set poisoning consistently causes substantial structural corruption, with mean per-digit ARI falling to 0.14–0.26. The effect also reproduces across ten independently published text encoders and is not confined to idealized embedding-space injection: text-derived poison passed through the same tokenizer-fitting pipeline reaches a comparable corruption regime. However, tokenizer corruption does not translate monotonically into recommendation harm. Across nine recommender–catalog settings trained to convergence, only three show significant downstream changes, and their directions differ across architectures, while substantial structural corruption remains present throughout. Matched benign-growth controls further show that ordinary catalog additions can induce structural drift comparable to malicious insertion, so drift magnitude alone is not a reliable attack signature. Six published anomaly-detection baselines fail to restore the clean tokenizer structure. A fit-aware trusted-reference admission gate strongly separates the non-adaptive attacks from held-out legitimate additions across three catalogs, but an adaptive adjacent-mimicry attack exposes an evasion region. Frozen-rotation warm-started refitting complements screening by constraining admitted updates, maintaining ARI at or above 0.98 across tested single-refit attack severities and geometries, while sequential drip-feed updates reveal gradual erosion under repeated refitting. Code will be provided after acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.