acceptodds
Under review as a conference paper at ICLR 2027

Pharos: From Safety Verdict to Routing Primitive at Inline Latency

Abstract

Safety filters often report whether a user prompt is harmful, but many applications also need to know which risk it presents. We study this finer-grained task under the 16 risk categories in the AIR 2024 Level-2 taxonomy. Risk Atlas unifies 23 public safety datasets under that taxonomy, preserves source provenance, and adds category-checked synthetic prompts to broaden coverage. Splits are stratified by source and risk category, with duplicate queries removed and newly synthesized prompts confined to training. On 10,000 prompts with independently mapped AIR labels, the labeling model reaches 0.97 macro-F1. Pharos, a compact encoder trained on Risk Atlas, reaches 0.753 AIR Level-2 macro-F1 on the held-out test set, ahead of evaluated released guards and a decoder trained on the same corpus. On an RTX 4080-SUPER, Pharos's median batch-one inference time is 5.36 ms per query. Together, Risk Atlas and Pharos provide a shared evaluation setting and a fast baseline for fine-grained prompt risk categorization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.