In Java, a security control is something you unit-test at a clean boundary — validate this string, reject that regex — and once it passes review it doesn't quietly get weaker because the caller phrased the same request differently. A prompt-injection red-team harness breaks that: there is no fixed input format to validate, only whatever natural-language string reaches the model, so "the same attack" has an unbounded number of equivalent phrasings, and a harness that only recognizes a handful of them is dishonest about what it actually covers. Prompt injection walked the attack; this page walks the harness that scores a model against many of them at once: probe → model → detector → report.
Every number below is a real run of exercises/sec-prompt-injection's reference
harness against its own fixture and its provided mock model. The model is deliberately naive:
it "complies" — and leaks a hidden token, canary-7f3a9c — only when a prompt
contains one of five literal trigger phrases, case-insensitively; anything else it refuses.
Across the fixture that gives 12 of 40 successful attacks (30.0%): 4 of 10
override, 5 of 10 exfiltration, 3 of 10 roleplay — and
0 of 10 encoding, because those ten are obfuscated (base64, reversed,
ROT13) and never contain a trigger phrase literally. Pick a probe below to walk one attack
through the whole pipeline.
detect_canary_leak only inspects the model's output for the token — it
never sees which phrase, if any, caused the model to comply. That's what makes it a legitimate
black-box detector: point the same function at a real model instead of the mock one and it
still works, because it never depended on knowing the model's weaknesses in advance. The
price of that generality shows up in atk-31: a base64-encoded instruction never
contains a trigger phrase as literal text, so this particular mock model — which only does a
substring match — never "sees" the attack and simply refuses. That is not the detector working
correctly; it's the model resisting an attack it was never actually exposed to. A
harness that reports 0% on the encoding row without saying so would let a reader
conclude obfuscated attacks don't work, when the honest reading is narrower: they don't work
against a keyword-only trigger, which is exactly why the report breaks the rate down
by category instead of only printing one number.