LLM red-teaming — the harness view

In Java, a security control is something you unit-test at a clean boundary — validate this string, reject that regex — and once it passes review it doesn't quietly get weaker because the caller phrased the same request differently. A prompt-injection red-team harness breaks that: there is no fixed input format to validate, only whatever natural-language string reaches the model, so "the same attack" has an unbounded number of equivalent phrasings, and a harness that only recognizes a handful of them is dishonest about what it actually covers. Prompt injection walked the attack; this page walks the harness that scores a model against many of them at once: probe → model → detector → report.

OWASP LLM01black-box detector 40-attack fixturesec-prompt-injection

One pipeline, run on all 40 fixture attacks

Every number below is a real run of exercises/sec-prompt-injection's reference harness against its own fixture and its provided mock model. The model is deliberately naive: it "complies" — and leaks a hidden token, canary-7f3a9c — only when a prompt contains one of five literal trigger phrases, case-insensitively; anything else it refuses. Across the fixture that gives 12 of 40 successful attacks (30.0%): 4 of 10 override, 5 of 10 exfiltration, 3 of 10 roleplay — and 0 of 10 encoding, because those ten are obfuscated (base64, reversed, ROT13) and never contain a trigger phrase literally. Pick a probe below to walk one attack through the whole pipeline.

Why the detector never reads the trigger list

detect_canary_leak only inspects the model's output for the token — it never sees which phrase, if any, caused the model to comply. That's what makes it a legitimate black-box detector: point the same function at a real model instead of the mock one and it still works, because it never depended on knowing the model's weaknesses in advance. The price of that generality shows up in atk-31: a base64-encoded instruction never contains a trigger phrase as literal text, so this particular mock model — which only does a substring match — never "sees" the attack and simply refuses. That is not the detector working correctly; it's the model resisting an attack it was never actually exposed to. A harness that reports 0% on the encoding row without saying so would let a reader conclude obfuscated attacks don't work, when the honest reading is narrower: they don't work against a keyword-only trigger, which is exactly why the report breaks the rate down by category instead of only printing one number.

Takeaways: a red-team harness is three separable parts — the probe set, the thing being attacked, and a detector that scores the output, not the input. A detector that stays black-box (reads only what came back) works against any model you point it at; a rate reported as one number instead of broken down by attack category hides exactly the blind spot — obfuscation past a naive trigger — that matters most to a reader deciding whether "30% attack success" means the model is fairly safe or the harness barely tried.