Do Detailed Guard Denials Help Bypass Guardrails? Evaluating Runtime Feedback in DeepSeek Harness
Abstract
Runtime guards block tool calls, but their denial messages re-enter an agent's context and can shape subsequent planning. Detailed diagnostics may help benign recovery while also providing actionable information to a planner already influenced by indirect prompt injection. Yet existing evaluations often change attack history, search strategy, and feedback together, making the effect of the denial itself difficult to isolate. We introduce GuardOracle, a controlled audit at the native denial boundary of DeepSeek Harness. Starting from the same pre-feedback history, GuardOracle varies only the model-visible denial representation—Opaque, a length-matched Placebo, Category, or Detailed—while holding the injection, guard policy, tools, simulator state, and post-denial action budget fixed. A bypass is counted only when a candidate action is permitted, committed in simulation, and independently verified to realize the frozen malicious objective. In our primary study, Detailed minus Opaque changes verified Bypass@3 by +7.85 percentage points (95% CI: −0.54 to 17.83), leaving the effect direction uncertain. In an independent prospective study, none of 658 frozen screening repetitions reached an eligible first-denial state. These results motivate separating whether an agent reaches the denial boundary (activation) from how feedback changes behavior once that boundary is reached (conditional feedback effect).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.