Skip to content

Detection benchmark ​

What the proxy's detection layer stops, measured on data it was not written against. The numbers on this page and in the README are generated from scripts/waf_eval/results.json; a test fails the build if they drift apart.

Results ​

Measured 2026-10-10 on llmproxy 1.38.1 (byte firewall and SecurityShield, default configuration), against protectai/deberta-v3-base-prompt-injection-v2 at threshold 0.5.

DatasetKindPromptsllmproxyClassifier
deepset/prompt-injectionsattack2634.6%41%
Lakera/gandalf_ignore_instructionsattack1,00039%100%
scripts/waf_eval/heldout.pyattack5227%90%
uiuc-kang-lab/InjecAgentattack1,0540%66%
jackhhao/jailbreak-classificationattack66632%84%
deepset/prompt-injectionsbenign3990%1.0%
scripts/waf_eval/heldout.pybenign424.8%38%
jackhhao/jailbreak-classificationbenign1,3320.1%0.6%
leolee99/NotInjectbenign3390%43%
All attacks, stopped3,03521%79%
All benign, stopped by mistake2,1120.1%8.2%

The held-out set by family (stopped / prompts):

KindFamilyllmproxyClassifier
attackdirect1 / 33 / 3
attackparaphrase0 / 66 / 6
attackextraction2 / 66 / 6
attackroleplay2 / 63 / 6
attackmultilingual4 / 88 / 8
attackobfuscation5 / 87 / 8
attackindirect0 / 54 / 5
attackexfiltration0 / 22 / 2
attackauthority0 / 33 / 3
attackhypothetical0 / 33 / 3
attackagent0 / 22 / 2
benignordinary0 / 61 / 6
benignsecurity-talk2 / 64 / 6
benigninstruction-words0 / 103 / 10
benigndevops0 / 82 / 8
benigndocument0 / 44 / 4
benignmultilingual0 / 41 / 4
benignfiction0 / 41 / 4

The tool policy on the indirect-injection benchmark (the model is assumed to obey the planted instruction; the policy is after_tool_result: ["*Get*", "*Read*", "*Search*", "*View*", "*List*", "*Navigate*"]):

Attack typeCasesAttacker's call refused
direct harm510510
data stealing544544
the user's own call, refused by mistake1,0540

The classifier at other thresholds (attacks stopped; hard benign prompts stopped by mistake):

ThresholdAttacksHard benign
0.579%42%
0.9963%29%
0.999950%16%

Hard benign prompts are the trigger-word set and the held-out negatives (381 prompts). On 1,054 clean tool results (the indirect-injection templates with the attacker's text replaced by an ordinary sentence) the classifier flagged 29%. Latency per prompt on a laptop CPU: llmproxy median 0.43 ms; classifier median 16 ms, 95th percentile 195 ms.

Reading the numbers ​

The proxy's detection is lexical. The byte firewall matches signatures after decoding common encodings; the shield scores regular expressions and compares character trigrams against a list of known phrasings. That recognises the well-known wordings and their obfuscations, in well under a millisecond and with almost no false positives. It does not recognise an attack that is worded differently, and it has nothing to say about an instruction planted in a tool result, which is an ordinary sentence in the wrong place.

A classifier is not the answer either. The open model used as a reference stops about four attacks in five, and refuses a large share of legitimate text that merely talks about instructions, as well as a good part of clean tool output. Raising its threshold trades the one for the other without a setting that is good at both. It is useful as a signal to record and review; as a gate it breaks ordinary work.

So do not rely on recognising the attack. What limits the damage of an injection is what the response is allowed to do: which tools may be called and when, where links and images may point, whether a secret can leave. Those controls do not depend on the wording of the attack.

Caveats ​

  • Three of the public datasets (gandalf, deepset, jackhhao) are widely used for training; the reference classifier has very likely seen them. The indirect-injection benchmark, the trigger-word set and the held-out set are the out-of-distribution tests.
  • The held-out set is small (94 prompts) and was written by the reviewer who ran the measurement, before reading the signatures. It is there so the repository carries a set the detector was not tuned on, not as a benchmark of its own.
  • "Stopped" means the request is refused. For the indirect-injection set the whole conversation (user turn, tool call, tool result) is sent as the proxy would receive it.
  • Each prompt is judged on its own: the shield's multi-turn and cross-session memory is empty.
  • The regression corpus in tests/corpus/ is not part of this: it was written with the detector.

Reproducing it ​

bash
make waf-bench            # fetch (about 750 MB), score, write results.json, update the tables
make waf-bench-check      # only the part that needs no download: the held-out set

scripts/waf_eval/fetch.sh downloads the datasets and the classifier at pinned revisions into .waf-bench/ (git-ignored); they are not redistributed here.

SourceLicence
deepset/prompt-injectionsApache-2.0
Lakera/gandalf_ignore_instructionsMIT
jackhhao/jailbreak-classificationApache-2.0
leolee99/NotInjectMIT
uiuc-kang-lab/InjecAgentMIT
protectai/deberta-v3-base-prompt-injection-v2 (reference classifier)Apache-2.0

MIT License