acceptodds
Under review as a conference paper at ICLR 2027

Cooldown Landmines: Cross-Tenant Interference Attacks on LLM Gateways

Abstract

LLM gateways enforce separate tenant quotas while sharing model deployments and cooldown records that temporarily exclude failing backends. However, a tenant's request failure can update these shared records and restrict other tenants' access to serviceable deployments. We identify two attacks that exploit this gap in LiteLLM. The first uses requests rejected at the key's requests-per-minute (RPM) limit: caller-supplied identifiers for known registered deployments reach failure handling, allowing two rejected requests to redirect another tenant to fallback with zero upstream calls from those requests. The second uses admitted traffic to create cooldown records that persist after backend capacity recovers. To address these failures, we design an origin check that blocks deployment updates from key RPM rejections and tenant-scoped cooldown that preserves the triggering tenant's back-off while retaining shared records for backend faults. Experiments with authenticated proxies and self-hosted vLLM demonstrate that the attacks can force fallback or denial while deployments remain serviceable. Across five paired four-worker runs, tenant scoping reduces victim fallback after recovery from to , while increasing attempts against exhausted shared quotas. These findings show that tenant isolation must cover both request admission and the failure handling that governs shared deployment availability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.