AI Defense Lab
Your AI Security Learning Platform
Start with one experiment
Same input. Different policy. What changes?
Open the Threat Lab, run its ready-made example, and inspect what the detectors noticed. Then try Monitor Only to see why detecting an attack and blocking it are different decisions.
Use the supplied examples: detection and policy evaluation run for real, while dangerous tool actions are simulated. Do not paste passwords, personal data, or confidential documents. An allowed result does not prove an input is safe.
What You'll Learn
Explore how attacks influence AI systems, what pattern-based detectors notice, and how policies respond. Follow the evidence, including what these defenses miss.
Learning Modules
Threat Detection Lab
Run attack payloads through the live detection pipeline and read the evidence it produces
Policy Design Lab
Build rules that decide what to block, allow, or flag for review
RAG Security Lab
Understand how retrieval-augmented generation can be poisoned and defended
Agent Attack Lab
See how multi-stage AI agents can be attacked at each pipeline stage
Command Injection Lab
See how model output becomes the injection payload for whatever runs next
Audit Trail Explorer
Read what a run recorded, and what that record is and is not evidence of
Lab Activity
These totals are shared across visitors to this lab, not your personal learning progress. Block rate is the share of policy-evaluated runs that were blocked; RAG retrieval queries are excluded because no policy judges them. It is not detector accuracy: the submitted inputs have no labelled ground truth against which to measure misses.
—
Total Runs
—
Runs Blocked
—
Active Policies
—
Block Rate
What these detectors actually catch
Measured against a labelled corpus of 34 attacks and 30 benign inputs, written from the threat model rather than from the detectors’ own pattern lists — so the test and the code cannot agree by construction. Regenerate with pytest tests/test_detection_efficacy.py -s -k report. Measured on 2026-10-09 on main commit f3f0a12 plus the change that regenerated this file.
Intervals are 95% Wilson score intervals on the counts above: how much a count this small can say, not whether the cases were well chosen. 0 cases errored (a detector raised); errors are counted as neither caught nor blocked and stay in the denominator.
Read this as a regression baseline, not a benchmark. The corpus was assembled by the same people who wrote the detectors, so 82% here is not comparable to a published evaluation on an independent adversarial dataset. It answers “has detection got worse?”, not “how would this fare against a real attacker?” For scale: At a 1% false-positive rate, five previously published detectors caught between 1.97% and 20.37% of injections on the PromptShield evaluation split, and between 0% and 9.39% at 0.1%. The ledger entry states the conditions behind that range, and why it is not comparable to this one (ACM CODASPY 2025 (extended technical report)).
By attack class
By whether patterns can reach it
Character-level obfuscation — homoglyphs, zero-width characters, spacing, leetspeak. Folding the input to a canonical form reaches these.
The hostile wording is plain English. A pattern can match it directly, so a miss here is missing coverage.
Paraphrase, other languages, or intent split across turns. No pattern list closes these; they need a model that understands meaning.
OWASP Top 10 for LLM Applications
2026 revision — verified 2026-10-07 against genai.owasp.org. Every identifier below links to its own source page.
“Detectors” below means one or more regex detectors in this lab emit that identifier. It is not a claim that the risk is handled. At a 1% false-positive rate, five previously published detectors caught between 1.97% and 20.37% of injections on the PromptShield evaluation split, and between 0% and 9.39% at 0.1%. The authors built that split and calibrated the thresholds on it, their training data are English only, and multi-turn and function-calling attacks were out of scope. One of the same detectors scores 71% on its vendor's own data where this study measured 12.8%: a detection rate depends on the data and the operating point. At least two of the five are learned classifiers, not pattern matchers. This lab's 82% on its own 34-case corpus is a third, different measurement and is not comparable. ACM CODASPY 2025 (extended technical report) (independent study, not a vendor claim).
| ID | Name | Description | Detectors |
|---|---|---|---|
| LLM01 | Prompt Injection | User input alters the model's behaviour or output in ways the developer did not intend. | override_detector, encoding_detector, poisoning_detector |
| LLM02 | Sensitive Information Disclosure | The model or its surrounding application exposes PII, credentials, business data, or its own training material. | exfil_detector |
| LLM03 | Excessive Agency | The model is granted more functionality, permission, or autonomy than the task requires. | tool_detector |
| LLM04 | Supply Chain | Compromised models, datasets, adapters, or packages enter the application through its dependencies. | No detector |
| LLM05 | Data and Model Poisoning | Training or fine-tuning data is manipulated to implant bias, backdoors, or degraded behaviour in the model itself. | No detector |
| LLM06 | Unbounded Consumption | Uncontrolled inference volume or size drives denial of service, cost exhaustion, or model extraction. | No detector |
| LLM07 | Misinformation | The model produces confident, plausible, and wrong output that downstream users or systems act on. | No detector |
| LLM08 | Hidden Context Exposure | Hidden context — the system prompt and the rules, tools and data placed beside it — is disclosed, revealing instructions, logic, or secrets. | override_detector |
| LLM09 | Vector and Embedding Weaknesses | The retrieval layer itself is the attack surface — poisoned documents, cross-tenant leakage, or inversion of stored embeddings. | poisoning_detector |
| LLM10 | Improper Output Handling | Model output is passed to a downstream system without validation, turning the LLM into an injection relay. | No detector |